Trained Model Generation Method and Trained Model Generation Device

By generating a learned model through the combination of multiple base and target models learned from different information sets, the method improves inference accuracy and reduces workload in robot object retrieval systems.

JP7693010B2Active Publication Date: 2025-06-16KYOCERA CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023548510
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-17
Filing Date
2022-09-15
Publication Date
2025-06-16
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

Existing systems for generating learned models for robot object retrieval are inefficient and require significant workload, as they typically rely on a single learned model without effectively utilizing multiple models with different parts.

Method used

The method involves generating a learned model by acquiring multiple base models with a first part and target models with a second part, learned based on different sets of information, and connecting these models to improve inference accuracy and reduce workload.

Benefits of technology

This approach enhances the inference accuracy of the learned model by leveraging multiple models and reduces the workload in generating the learned model, allowing for more efficient recognition of objects in input information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693010000001
    Figure 0007693010000001
  • Figure 0007693010000002
    Figure 0007693010000002
  • Figure 0007693010000003
    Figure 0007693010000003
Patent Text Reader

Abstract

This trained model generation method involves generating a trained model that outputs, on the basis of a plurality of models each having at least one of a first part and a second part, a recognition result of a recognition target included in input information. In generating the trained model, a plurality of base models each having a portion corresponding to the first part are acquired, the base models having been trained on the basis of at least one set of first information related to the input information. In generating the trained model, target models each having a portion corresponding to the second part are acquired, the target models having been trained on the basis of at least one set of second information related to the input information in a state of being connected to the plurality of base models. In generating the trained model, generated is a trained model which has a target model having at least a portion corresponding to the second part.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims the priority of Japanese Patent Application No. 2021-152586 (filed on September 17, 2021), and the entire disclosure of the application is incorporated herein by reference for that purpose.

Technical Field

[0002] This disclosure relates to a method for generating a learned model, an inference device, and a learned model generation device.

Background Art

[0003] Conventionally, a system for causing a robot to take out an object using a learned model has been known (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

[0005] A method for generating a learned model according to an embodiment of the present disclosure includes generating a learned model that outputs a recognition result of a recognition target included in input information based on a plurality of models each having at least one of a first part and a second part. In the generation of the learned model, a plurality of base models each having a part corresponding to the first part, which are learned based on at least one set of first information related to the input information, are acquired. In the generation of the learned model, a target model each having a part corresponding to the second part, which is learned based on at least one set of second information related to the input information in a state of being connected to the plurality of base models, is acquired. In the generation of the learned model, a learned model having a target model having at least a part corresponding to the second part is generated.

[0006] An inference device according to an embodiment of the present disclosure includes a learned model that is generated based on a plurality of models each having at least one of a first part and a second part, and outputs a recognition result of a recognition target included in input information. The learned model has at least a target model having a part corresponding to the second part. The target model having a part corresponding to each of the second parts is a model obtained by learning based on at least one set of second information related to the input information while being connected to a plurality of base models each having a part corresponding to the first part. The plurality of base models each having a part corresponding to the first part are models learned based on at least one set of first information related to the input information.

[0007] A learned model generation device according to an embodiment of the present disclosure includes a control unit that generates a learned model that outputs a recognition result of a recognition target included in input information based on a plurality of models each having at least one of a first part and a second part. In generating the learned model, the control unit acquires a plurality of base models each having a part corresponding to the first part, which are learned based on at least one set of first information related to the input information. The control unit acquires a target model having a part corresponding to the second part, which is learned based on at least one set of second information related to the input information while being connected to the plurality of base models. The control unit generates a learned model having at least the target model having a part corresponding to the second part.

Brief Description of Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0009] (Configuration Example of Learned Model Generation System 1) The learned model generation system 1 according to an embodiment of the present disclosure generates a learned model 50 (see FIG. 2 etc.) that outputs a recognition result of an object to be recognized included in input information. The learned model generation system 1 generates a plurality of preliminary models as preparations for generating the learned model 50, updates the models by learning the models in which a part of each preliminary model is connected, and generates the learned model 50. The learned model generation system 1 can improve the inference accuracy of the learned model 50 by learning the models in which a plurality of preliminary models are connected. In addition, the learned model generation system 1 can reduce the workload for generating the learned model 50 by generating the learned model 50 using a part of the preliminary models. Note that, in this specification, the preliminary model is also referred to as a base model. The model generated by learning the model in which the preliminary models are connected is also referred to as a target model.

[0010] As shown in FIG. 1, the learned model generation system 1 according to an embodiment of the present disclosure includes a preliminary model generation device 10 and a learned model generation device 20. The learned model generation system 1 generates a preliminary model by the preliminary model generation device 10 and generates a learned model 50 by the learned model generation device 20. The preliminary model generation device 10 and the learned model generation device 20 may be configured as separate devices or as an integrated device.

[0011] The preliminary model generation device 10 includes a first control unit 12 and a first interface 14. The learned model generation device 20 includes a second control unit 22 and a second interface 24. The descriptions of "first" and "second" are simply given to distinguish the configurations included in different devices respectively. The first control unit 12 and the second control unit 22 are also simply referred to as control units. The first interface 14 and the second interface 24 are also simply referred to as interfaces.

[0012] <Control Unit> The control unit may be configured to include at least one processor to provide control and processing capabilities for executing various functions. The processor may execute a program for realizing various functions of the control unit. The processor may be realized as a single integrated circuit. The integrated circuit is also referred to as an IC (Integrated Circuit). The processor may be realized as a plurality of communicably connected integrated circuits and discrete circuits. The processor may be realized based on various other known technologies.

[0013] The control unit may include a storage unit. The storage unit may include an electromagnetic storage medium such as a magnetic disk, or may include a memory such as a semiconductor memory or a magnetic memory. The storage unit stores various information. The storage unit stores programs and the like executed by the control unit. The storage unit may be configured as a non-temporary readable medium. The storage unit may function as a work memory of the control unit. At least a part of the storage unit may be configured separately from the control unit.

[0014] <Interface> The first interface 14 of the preliminary model generation device 10 and the second interface 24 of the learned model generation device 20 input and output information or data to and from each other. The first interface 14 outputs the information or data acquired from the first control unit 12 to the learned model generation device 20, and outputs the information or data acquired from the learned model generation device 20 to the first control unit 12. The second interface 24 outputs the information or data acquired from the preliminary model generation device 10 to the second control unit 22. The interface may be configured to include a communication device that can communicate wired or wirelessly. The interface is also referred to as a communication unit. The communication device may be configured to be able to communicate by a communication method based on various communication standards. The interface can be configured by known communication technologies.

[0015] (Configuration example of the learned model 50) As shown in FIG. 2, the learned model 50 is represented as a model that connects a first learned model 51 and a second learned model 52. The first learned model 51 is also referred to as a backbone. The second learned model 52 is also referred to as a head. The backbone is configured to output the result of extracting the feature amount of the input information. The feature amount represents, as a numerical value, the feature of the appearance such as, for example, the edge or pattern of the learning target. The backbone may include, for example, convolution and pooling. The head is configured to make a predetermined determination about the input information based on the output of the backbone. Specifically, the head may output the recognition result of the recognition target included in the input information based on the extraction result of the feature amount of the input information output by the backbone. The head may include a fully connected layer that processes the extraction result of the feature amount by the backbone. Note that the head may include convolution and pooling. The head is configured to execute the recognition of the recognition target as a predetermined determination. For example, in the task of distinguishing a horse from a zebra, the feature amount may be a parameter representing the ratio of the area of the stripe pattern on the body surface. The predetermined determination may be to compare the ratio of the area of the stripe pattern on the body surface with a threshold value and determine whether the recognition target is a horse or a zebra. Also, for example, in the task of distinguishing an abalone from a turban shell, the feature amount may be a parameter representing the size or the number of holes in the shell. The predetermined determination may be to compare the size or the number of holes in the shell with a threshold value and determine whether the recognition target is an abalone or a turban shell.

[0016] The learned model 50 may be configured to include a CNN (Convolution Neural Network) having a plurality of layers. For the information input to the learned model 50, convolution based on a predetermined weighting coefficient is executed in each layer of the CNN. In the learning of the learned model 50, the weighting coefficient is updated. The learned model 50 may be configured to include a fully connected layer. The learned model 50 may be configured by VGG16 or ResNet50. The learned model 50 may be configured as a transformer. The learned model 50 is not limited to these examples and may be configured as various other models.

[0017] (Operation Example of Trained Model Generation System 1) In the trained model generation system 1, the trained model generation device 20 pre-generates or acquires a plurality of preliminary models including a backbone and a head. The trained model generation device 20 prepares one learning head in order to generate the head part of the trained model 50. The trained model generation device 20 sequentially connects the backbone of each of the plurality of preliminary models to one learning head. The trained model generation device 20 performs learning on the model obtained by connecting the backbone of each preliminary model and the learning head, and updates the learning head. The trained model generation device 20 sequentially connects the backbone of each preliminary model to the learning head, performs learning on each model, and updates the learning head. When the learning on the model obtained by connecting the backbone of each preliminary model is completed, the trained model generation device 20 applies the trained head as the head of the trained model 50. Further, the trained model generation device 20 separately generates or acquires the backbone of the trained model 50. The trained model generation device 20 generates the trained model 50 by connecting the separately generated or acquired trained head to the backbone.

[0018] The trained model generation device 20 may generate the backbone part of the trained model 50 by performing learning on the model obtained by sequentially connecting the head parts of each preliminary model. In the trained model generation device 20, the trained model 50 may be generated by sequentially connecting the preliminary model to the learning model and performing learning on both the backbone part and the head part of the trained model 50.

[0019] The learned model generation device 20 can be said to generate the learned model 50 by updating the pre-learning model through learning. The pre-learning model is a model obtained by connecting a first pre-learning model corresponding to the first pre-learned preliminary model 41 and a second pre-learning model corresponding to the second pre-learned preliminary model 42. The learned model generation device 20 may update the second pre-learning model to generate the second learned model 52 by learning in a model in which the first pre-learned preliminary model 41 is connected to the second pre-learning model instead of the first pre-learning model. The learned model generation device 20 may update the first pre-learning model to generate the first learned model 51 by learning in a model in which the second pre-learned preliminary model 42 is connected to the first pre-learning model instead of the second pre-learning model.

[0020] In the present embodiment, the model is configured to have at least one of a first part and a second part. That is, the base model is configured to have at least one of a first part and a second part. The first pre-learning model of the pre-learning model corresponds to the first part of the model. The second pre-learning model of the pre-learning model corresponds to the second part of the model. The first learned model 51 of the learned model 50 corresponds to the first part of the model. The second learned model 52 of the learned model 50 corresponds to the second part of the model.

[0021] <Generation of a head based on learning in a model connecting a plurality of backbones> Hereinafter, in the learned model generation system 1, an operation example of generating the head of the learned model 50 by learning a model in which the backbone of the preliminary model is connected to the learning head will be described.

[0022] As shown in FIG. 3, the first control unit 12 of the preliminary model generation device 10 generates or acquires a plurality of pre-learning preliminary models 301 to 30N in advance. The pre-learning preliminary models 301 to 30N are also collectively referred to as the pre-learning preliminary model 30.

[0023] The pre-learning preliminary model 301 includes a first pre-learning preliminary model 311 and a second pre-learning preliminary model 321. The pre-learning preliminary model 30N includes a first pre-learning preliminary model 31N and a second pre-learning preliminary model 32N. The first pre-learning preliminary models 311 to 31N of the pre-learning preliminary model 30 correspond to the first part of the model. The second pre-learning preliminary models 321 to 32N of the pre-learning preliminary model 30 correspond to the second part of the model. In FIG. 3, it is assumed that the configuration of the CNN layers or the filter size of each layer in the first pre-learning preliminary model 311 is different from the configuration of the CNN layers or the filter size of each layer in the first pre-learning preliminary model 31N. That is, it is assumed that the configuration of the CNN layers or the filter size of each layer in each of the first pre-learning preliminary models 311 to 31N is different from each other. The filter size means the size of the filter used to perform convolution (downsampling) or transposed convolution (upsampling) in the CNN. The model configurations of each of the first pre-learning preliminary models 311 to 31N may be different from each other or the same. Also, in FIG. 3, it is assumed that the configuration of the fully connected layers or the parameter size of each layer in the second pre-learning preliminary model 321 is the same as the configuration of the fully connected layers or the parameter size of each layer in the second pre-learning preliminary model 32N. That is, it is assumed that the configuration of the fully connected layers or the parameter size of each layer in each of the second pre-learning preliminary models 321 to 32N is the same. The configuration of the fully connected layers or the parameter size of each layer in each of the second pre-learning preliminary models 321 to 32N may be different from each other. The parameter size refers to the number of units constituting the fully connected layer, etc. The model configurations of each of the second pre-learning preliminary models 321 to 32N may be different from each other or the same.

[0024] The first control unit 12 learns each pre-training model 30 using, as learning data, first information that is the same as or related to the input information input to the learned model 50. The first information may be configured as a set of a plurality of learning images. The first control unit 12 may learn each pre-training model 30 using the same set of first information as learning data, or may learn each pre-training model 30 using different sets of first information as learning data. That is, the first control unit 12 may execute learning of the pre-training model 30 based on at least one set of first information. The first control unit 12 updates each pre-training model 30 by learning, and generates a plurality of learned pre-training models 401 to 40N. The learned pre-training models 401 to 40N are also collectively referred to as the learned pre-training model 40. The learning data may include teacher data used in so-called supervised learning. The learning data may include data generated by the device itself that performs learning, which is used in so-called unsupervised learning.

[0025] The learned pre-training model 401 includes a first learned pre-training model 411 and a second learned pre-training model 421. The learned pre-training model 40N includes a first learned pre-training model 41N and a second learned pre-training model 42N. The first learned pre-training models 411 to 41N of the learned pre-training model 40 correspond to the first part of the base model. The second learned pre-training models 421 to 42N of the learned pre-training model 40 correspond to the second part of the base model. The configuration of the CNN layers or the filter size of each layer in the first learned pre-training models 411 to 41N is the same as the configuration of the CNN layers or the filter size of each layer in the first pre-training models 311 to 31N. The configuration of the fully connected layers or the parameter size of each layer in the second learned pre-training models 421 to 42N is the same as the configuration of the fully connected layers or the parameter size of each layer in the second pre-training models 321 to 32N.

[0026] The second control unit 22 of the learned model generation device 20 acquires a learned preliminary model 40 from the preliminary model generation device 10 as a preliminary model. The preliminary model generation device 10 may output the learned preliminary model 40 to the learned model generation device 20 via the first interface 14. The learned model generation device 20 may acquire the learned preliminary model 40 from the preliminary model generation device 10 via the second interface 24.

[0027] The second control unit 22 performs learning on a model in which the backbone of each preliminary model and the learning head are connected, using second information that is the same as or related to the input information input to the learned model 50 as learning data. The backbone of each preliminary model corresponds to the first part of the base model. The learning head corresponds to the second part of the target model. The second information may be the same as or different from the first information. The second information may be configured as a set of a plurality of learning images. The second control unit 22 may perform learning on each model in which the backbone of each preliminary model and the learning head are connected, using the same set of second information as learning data, or may perform learning using different sets of second information as learning data. Further, the second control unit 22 may divide one set of second information into even smaller subsets and perform learning using different subsets as learning data each time the backbone connected to the learning head is changed. The second control unit 22 may perform learning using the same subset as learning data when the backbone connected to the learning head is changed. That is, the second control unit 22 may perform learning on a model in which the backbone of each preliminary model and the learning head are connected based on at least one set of second information. Note that the amount of information of the second information used for learning the learned model 50 may be equal to or less than the amount of information of the first information used for learning the preliminary model. Note that the amount of information refers to, for example, the number of learning images included in the second information.

[0028] Specifically, as shown in FIG. 4, the second control unit 22 sequentially performs learning on a model obtained by transferring the backbone of each preliminary model (the first learned preliminary models 411 to 41N) and connecting it to the second pre-learning model 520 or the second in-learning models 521 to 52(N - 1). The second control unit 22 may perform learning according to the following procedure. The second control unit 22 performs learning on a model obtained by connecting the first learned preliminary model 411 and the second pre-learning model 520, and updates the second pre-learning model 520 to the second in-learning model 521. The second control unit 22 performs learning on a model obtained by connecting the first learned preliminary model 412 and the second in-learning model 521 updated in the previous step, and updates the second in-learning model 521 to the second in-learning model 522. The second control unit 22 performs learning on a model obtained by connecting the first learned preliminary model 41N and the second in-learning model 52(N - 1) updated in the previous step, and updates the second in-learning model 52(N - 1) to the second learned model 52N.

[0029] By executing the procedure described above, the second control unit 22 generates the second learned model 52N. The second control unit 22 applies the second learned model 52N as the second learned model 52 (the head of the learned model 50). The second learned model 52N generated by connecting and learning to the first part of the base model corresponds to the second part of the target model.

[0030] <Generation of Backbone Based on Learning in a Model with Multiple Connected Heads> Hereinafter, in the learned model generation system 1, an operation example of generating the backbone of the learned model 50 by performing learning on a model obtained by connecting the head of the preliminary model to the backbone for learning will be described.

[0031] As shown in FIG. 5, the first control unit 12 of the preliminary model generation device 10 generates or acquires a plurality of pre-learning preliminary models 301 to 30N in advance. In FIG. 5, it is assumed that the configuration of the fully connected layer or the parameter size of each layer in the second pre-learning preliminary model 321 is different from the configuration of the fully connected layer or the parameter size of each layer in the second pre-learning preliminary model 32N. That is, it is assumed that the configuration of the fully connected layer or the parameter size of each layer in each of the second pre-learning preliminary models 321 to 32N is different from each other. Also, in FIG. 5, it is assumed that the configuration of the CNN layer or the filter size of each layer in the first pre-learning preliminary model 311 is the same as the configuration of the CNN layer or the filter size of each layer in the first pre-learning preliminary model 31N. That is, it is assumed that the configuration of the CNN layer or the filter size of each layer in each of the first pre-learning preliminary models 311 to 31N is the same. The configuration of the CNN layer or the filter size of each layer in each of the first pre-learning preliminary models 311 to 31N may be different from each other.

[0032] The first control unit 12 learns each pre-learning preliminary model 30 using the first information that is the same as or related to the input information input to the learned model 50 as learning data. The first control unit 12 updates each pre-learning preliminary model 30 by learning and generates a plurality of learned preliminary models 401 to 40N. The configuration of the fully connected or CNN layer, the parameter size of each layer, or the filter size of each learned preliminary model 40 is the same as the configuration of the fully connected or CNN layer, the parameter size of each layer, or the filter size of each pre-learning preliminary model 30.

[0033] The second control unit 22 of the learned model generation device 20 acquires the learned preliminary model 40 from the preliminary model generation device 10 as a preliminary model. The second control unit 22 uses, as learning data, third information that is the same as or related to the input information input to the learned model 50, and executes learning on a model in which the head of each preliminary model is connected to a learning backbone. The head of each preliminary model corresponds to the second part of the base model. The learning backbone corresponds to the first part of the target model. The third information may be the same as or different from the first information or the second information. The third information may be configured as a set of a plurality of learning images. The second control unit 22 may perform learning on each model in which the head of each preliminary model is connected to a learning backbone, using the same set of third information as learning data, or may perform learning using different sets of third information as learning data. Further, the second control unit 22 may divide one set of third information into even smaller subsets, and perform learning using different subsets as learning data each time the head connected to the learning backbone is changed. The second control unit 22 may perform learning using the same subset as learning data when the head connected to the learning backbone is changed. That is, the second control unit 22 may execute learning on a model in which the head of each preliminary model is connected to a learning backbone based on at least one set of third information. Note that the amount of information of the third information used for learning the learned model 50 may be equal to or less than the amount of information of the first information used for learning the preliminary model. Note that the amount of information refers to, for example, the number of learning images included in the second information.

[0034] Specifically, as shown in FIG. 6, the second control unit 22 sequentially performs learning on a model obtained by transferring the heads of the respective preliminary models (the second learned preliminary models 421 to 42N) and connecting them to the first pre-learning model 510 or the first in-learning models 511 to 51(N - 1). The second control unit 22 may perform learning according to the following procedure. The second control unit 22 performs learning on a model obtained by connecting the second learned preliminary model 421 and the first pre-learning model 510, and updates the first pre-learning model 510 to the first in-learning model 511. The second control unit 22 performs learning on a model obtained by connecting the second learned preliminary model 422 and the first in-learning model 511 updated in the previous step, and updates the first in-learning model 511 to the first in-learning model 512. The second control unit 22 performs learning on a model obtained by connecting the second learned preliminary model 42N and the first in-learning model 51(N - 1) updated in the previous step, and updates the first in-learning model 51(N - 1) to the first learned model 51N.

[0035] By executing the procedure described above, the second control unit 22 generates the first learned model 51N. The second control unit 22 applies the first learned model 51N as the first learned model 51 (the backbone of the learned model 50). The first learned model 51N generated by connecting and learning to the second part of the base model corresponds to the first part of the target model.

[0036] <Generation of the learned model 50> The second control unit 22 of the learned model generation device 20 generates the learned model 50 by connecting the first learned model 51 and the second learned model 52.

[0037] When the second control unit 22 generates the second learned model 52 (head) based on the preliminary model, the second control unit 22 connects the first learned model 51 (backbone) to the generated second learned model 52 to generate the learned model 50. The second control unit 22 may generate the first learned model 51 by other means or obtain it from another device. The second control unit 22 may obtain at least one of the plurality of first learned preliminary models 41 as the first learned model 51.

[0038] When the second control unit 22 generates the first learned model 51 (backbone) based on the preliminary model, the second control unit 22 connects the second learned model 52 (head) to the generated first learned model 51 to generate the learned model 50. The second control unit 22 may generate the second learned model 52 by other means or obtain it from another device. The second control unit 22 may obtain at least one of the plurality of second learned preliminary models 42 as the second learned model 52.

[0039] The second control unit 22 may generate both the first learned model 51 (backbone) and the second learned model 52 (head) based on the preliminary model. The second control unit 22 connects the first learned model 51 and the second learned model 52 generated based on the preliminary model to generate the learned model 50.

[0040] The second control unit 22 may generate the learned model 50 that includes only the second learned model 52 generated by connecting to and learning from the first learned preliminary model 41. The second control unit 22 may generate the learned model 50 that includes only the first learned model 51 generated by connecting to and learning from the second learned preliminary model 42.

[0041] <parentheses> The learned model generation system 1 generates a preliminary model with the preliminary model generation device 10 and generates a learned model 50 based on the preliminary model with the learned model generation device 20. As shown in FIGS. 7 and 8, the learned model generation system 1 generates a learned preliminary model 40 including a first learned preliminary model 41 and a second learned preliminary model 42 by having the preliminary model generation device 10 learn the first information as learning data.

[0042] As shown in FIG. 7, the learned model generation system 1 may transfer the first learned preliminary model 41 to the learned model generation device 20. The learned model generation device 20 may generate a second in-learning model 521 to 52(N - 1) or a second learned model 52N by learning the second information as learning data in a model in which each of the first learned preliminary models 411 to 41N is connected to the second pre-learning model 520 or the second in-learning models 521 to 52(N - 1).

[0043] As shown in FIG. 8, the learned model generation system 1 may transfer the second learned preliminary model 42 to the learned model generation device 20. The learned model generation device 20 may generate a first in-learning model 511 to 51(N - 1) or a first learned model 51N by learning the third information as learning data in a model in which each of the second learned preliminary models 421 to 42N is connected to the first pre-learning model 510 or the first in-learning models 511 to 51(N - 1).

[0044] The learned model generation device 20 may apply the generated second learned model 52N as the second learned model 52. The learned model generation device 20 may apply any one of the second in-learning models 521 to 52(N - 1) as the second learned model 52. The learned model generation device 20 may apply the generated first learned model 51N as the first learned model 51. The learned model generation device 20 may apply any one of the first in-learning models 511 to 51(N - 1) as the first learned model 51.

[0045] The trained model generation device 20 generates a trained model 50 by connecting the first trained model 51 to the generated second trained model 52. The trained model generation device 20 may select the first trained model 51 to be connected to the generated second trained model 52 from among a plurality of first trained preliminary models 41. The trained model generation device 20 may obtain the first trained model 51 to be connected to the generated second trained model 52 from an external device. The trained model generation device 20 generates a trained model 50 by connecting the generated second trained model 52 to the generated first trained model 51. The trained model generation device 20 may select the second trained model 52 to be connected to the generated first trained model 51 from among a plurality of second trained preliminary models 42. The trained model generation device 20 may obtain the second trained model 52 to be connected to the generated first trained model 51 from an external device. The trained model generation device 20 may generate a trained model 50 by connecting the generated first trained model 51 and the generated second trained model 52.

[0046] <Example of the procedure of the method for generating the trained model 50> The trained model generation device 20 may execute a method for generating the trained model 50 including the procedure of the flowchart illustrated in FIG. 9. The method for generating the trained model 50 may be realized as a program for generating the trained model 50 to be executed by a processor constituting the second control unit 22 of the trained model generation device 20. The program for generating the trained model 50 may be stored in a non-transitory computer-readable medium.

[0047] The second control unit 22 acquires a plurality of trained preliminary models 40 from the preliminary model generation device 10 (step S1). The second control unit 22 generates a model in which the first trained preliminary model 41 and the second pre-training model 520 of each trained preliminary model 40 are connected (step S2). The second control unit 22 updates the second pre-training model 520 by learning the model generated in step S2, and generates a second in-training model 521 (step S3).

[0048] The second control unit 22 determines whether all of the first learned preliminary models 41 of the plurality of learned preliminary models 40 are connected (step S4). When not all of the first learned preliminary models 41 are connected (step S4: NO), the second control unit 22 returns to the procedure of step S2 and generates a model in which the yet-to-be-connected first learned preliminary model 41 is connected to the second in-learning models 521 to 52(N - 1). Further, the second control unit 22 updates the second in-learning models 521 to 52(N - 1) in the procedure of step S3 and generates the second in-learning models 522 to 52(N - 1) or the second learned models 52N.

[0049] When all of the first learned preliminary models 41 are connected (step S4: YES), the second control unit 22 generates a model in which the second learned preliminary model 42 of each learned preliminary model 40 and the first pre-learning model 510 are connected (step S5). The second control unit 22 updates the first pre-learning model 510 by learning the model generated in step S5 and generates the first in-learning model 511 (step S6).

[0050] The second control unit 22 determines whether all of the second learned preliminary models 42 of the plurality of learned preliminary models 40 are connected (step S7). When not all of the second learned preliminary models 42 are connected (step S7: NO), the second control unit 22 returns to the procedure of step S5 and generates a model in which the yet-to-be-connected second learned preliminary model 42 is connected to the first in-learning models 511 to 51(N - 1). Further, the second control unit 22 updates the first in-learning models 511 to 51(N - 1) in the procedure of step S6 and generates the first in-learning models 512 to 51(N - 1) or the first learned models 51N.

[0051] When the second control unit 22 connects all the second pre-trained models 42 (step S7: YES), it connects the first pre-trained model 51 and the second pre-trained model 52 to generate a pre-trained model 50 (step S8). Specifically, the second control unit 22 applies the first pre-trained model 51N updated and generated in the procedures of steps S2 and S3 as the first pre-trained model 51. The second control unit 22 applies the second pre-trained model 52N updated and generated in the procedures of steps S5 and S6 as the second pre-trained model 52. After executing the procedure of step S8, the second control unit 22 ends the execution of the flowchart procedure in FIG. 9. After executing the procedure of step S4, the second control unit 22 may proceed to the procedure of step S8 without executing the procedures from step S5 to S7. After executing the procedure of step S1, the second control unit 22 may proceed to the procedure of step S5 without executing the procedures from step S2 to S4.

[0052] As described above, the pre-trained model generation system 1 and the pre-trained model generation device 20 according to this embodiment generate a plurality of preliminary models and generate a pre-trained model 50 using each preliminary model. The pre-trained model generation device 20 generates a model in which a part of a plurality of preliminary models is connected to a learning model corresponding to the first pre-trained model 51 or the second pre-trained model 52 that is a part of the pre-trained model 50. The pre-trained model generation device 20 learns the generated model and updates the learning model to generate the first pre-trained model 51 or the second pre-trained model 52. The pre-trained model generation device 20 generates a pre-trained model 50 using the generated first pre-trained model 51 or second pre-trained model 52. By learning the model connected to the plurality of preliminary models, the recognition accuracy in various recognition targets can be improved on average. As a result, the recognition accuracy in recognition using the pre-trained model 50 can be improved.

[0053] (Meaning of the same or related information) The trained model generation system 1 according to this embodiment uses, as learning data, information that is the same as or related to the input information input to the trained model 50. The information that is the same as or related to the input information may be information on a task that is the same as or related to the task executed by the trained model 50 that receives the input information. For example, when the task is the classification of mammals included in an image, an example of the input information is an image depicting an organism including a mammal. And the information regarding the learning target generated as information on the same task as the input information is an image of a mammal. Also, the information regarding the learning target generated as information on a task related to the input information is, for example, an image of a reptile.

[0054] The task may include, for example, a classification task of classifying at least two types of recognition targets included in the input information. The classification task can be subdivided into, for example, a task of distinguishing whether the recognition target is a dog or a cat, or a task of distinguishing whether the recognition target is a cow or a horse. The task is not limited to the classification task and may include tasks for realizing various other operations. The task may include segmentation for determining from pixels belonging to a specific object. The task may include object detection for detecting an enclosing rectangular region. The task may include pose estimation of an object. The task may include keypoint detection for finding a certain feature point.

[0055] Here, when both the input information and the information regarding the learning target are information on the classification task, it is assumed that the relationship between the input information and the information regarding the learning target is information on related tasks. Further, when both the input information and the information regarding the learning target are information on a task of distinguishing whether the recognition target is a dog or a cat, it is assumed that the relationship between the input information and the information regarding the learning target is information on the same task. The relationship between the input information and the information regarding the learning target is not limited to these examples and can be determined under various conditions.

[0056] (Other embodiments) Hereinafter, other embodiments will be described.

[0057] (When making the first information different from the second information or the third information) The second control unit 22 of the learned model generation device 20 may make the second information or the third information used as learning data for learning different from the first information used as learning data for learning to generate a preliminary model. For example, when the first information as learning data for learning to generate a preliminary model is information for recognizing industrial parts, the second control unit 22 may use, as the second information or the third information, information for recognizing fine types of screws specialized for only screws among the industrial parts. For example, when the first information as learning data for learning to generate a preliminary model is information for recognizing animals, the second control unit 22 may use, as the second information or the third information, information for recognizing fine types of dogs specialized for only dogs among the animals. A learned model 50 that recognizes in a broad classification such as industrial parts or animals is also referred to as a general-purpose model. A learned model 50 that recognizes in a narrow classification such as the type of screw or the type of dog is also referred to as a dedicated model.

[0058] The second control unit 22 may make the granularity of the second information or the third information smaller than the granularity of the first information. The granularity of information means the fineness of the classification of the recognition target. For example, assume that the learned model 50 recognizes industrial parts as the recognition target. The granularity of information for classifying industrial parts into screws, nuts, washers, brackets, etc. is larger than the granularity of information for classifying screws by length or diameter, etc. In other words, the granularity of information differs according to the broad classification, medium classification, or small classification of the recognition target. The smaller the granularity of the information used as learning data, the more finely the learned model 50 can recognize the differences of the recognition target. On the other hand, when the granularity of the information used as learning data is too small, the learned model 50 may not be able to recognize the large differences of the recognition target. For example, a learned model 50 that can recognize the difference in the length or diameter of a screw may not be able to recognize the difference between a screw and a nut. A learned model 50 generated by learning using information with a large granularity as learning data corresponds to a general-purpose model. A learned model 50 generated by learning using information with a small granularity as learning data corresponds to a dedicated model.

[0059] (Evaluation and Regeneration of Trained Model 50) The second control unit 22 of the trained model generation device 20 may evaluate the recognition accuracy of the recognition target by the generated trained model 50. The second control unit 22 may regenerate the trained model 50 based on the evaluation result of the recognition accuracy.

[0060] Specifically, the second control unit 22 acquires the recognition result output from the trained model 50 when input information is input to the generated trained model 50. The second control unit 22 may input information for which the correct recognition result is known as the input information to the trained model 50, and evaluate the ratio (correct answer rate) at which the obtained recognition result matches the correct recognition result. The second control unit 22 may calculate the correct answer rate as an evaluation value. In this case, the higher the evaluation value, the higher the recognition accuracy of the trained model 50. When the evaluation value is equal to or greater than a predetermined threshold, the second control unit 22 may determine that the recognition accuracy of the generated trained model 50 is sufficient. When the evaluation value is less than a predetermined threshold, the second control unit 22 may determine that the recognition accuracy of the generated trained model 50 is insufficient.

[0061] When the recognition accuracy of the trained model 50 is insufficient, the trained model generation system 1 may regenerate the trained model 50. When regenerating the trained model 50, the second control unit 22 may change the second information or the third information used as the learning data in the learning executed before regenerating the trained model 50.

[0062] On the other hand, before regenerating the learned model 50, the second control unit 22 may learn the same information as the second information or the third information used as learning data in the learning executed without changing them. In this case, the second control unit 22 may change the combination of the preliminary model connected to the learning model and the set of the second information or the third information from the combination in the learning data used in the learning executed before regenerating the learned model 50. Also, when the second control unit 22 divides the second information or the third information into small groups, the second control unit 22 may change the combination of the preliminary model connected to the learning model and the small groups of the second information or the third information from the combination in the learning data used in the learning executed before regenerating the learned model 50. The second control unit 22 may change the order of the information used as learning data with respect to the order in which the combination of the learning model and the preliminary model is changed. That is, the second control unit 22 may shuffle the order of the information used as learning data. The second control unit 22 may change the information used as learning data or change the order of the information used as learning data with respect to the combination of the learning model and the preliminary model until the recognition accuracy of the learned model 50 becomes equal to or higher than a predetermined accuracy, and regenerate the learned model 50.

[0063] Also, when the second control unit 22 learns the same information as the second information or the third information without changing them, the second control unit 22 may change the configuration of the small groups of the second information or the third information. That is, the second control unit 22 may change the content of the small groups while using the same set of learning data, and regenerate the learned model 50.

[0064] Note that the method for evaluating and regenerating the learned model 50 may also be applied to the learning using the first information.

[0065] In the trained model 50 generated as the target model, if there is a target with poor recognition accuracy, the second control unit 22 may generate the target model by relearning. In this case, the second control unit 22 may regenerate the trained model 50 as a new target model by learning based on new learning data without using a preliminary model (base model). Note that the new learning data is also referred to as fourth information. The fourth information may be information identical to or related to the input information.

[0066] (Configuration example of robot control system 100) As shown in FIG. 10, a robot control system 100 according to an embodiment includes a robot 2 and a robot control device 110. In this embodiment, it is assumed that the robot 2 moves the work object 8 from the work start point 6 to the work target point 7. That is, the robot control device 110 controls the robot 2 so that the work object 8 moves from the work start point 6 to the work target point 7. The work object 8 is also referred to as a work target. The robot control device 110 controls the robot 2 based on information regarding the space in which the robot 2 performs work. The information regarding the space is also referred to as space information.

[0067] <Robot 2> The robot 2 includes an arm 2A and an end effector 2B. The arm 2A may be configured as, for example, a 6-axis or 7-axis vertical articulated robot. The arm 2A may be configured as a 3-axis or 4-axis horizontal articulated robot or a scalar robot. The arm 2A may be configured as a 2-axis or 3-axis orthogonal robot. The arm 2A may be configured as a parallel link robot or the like. The number of axes constituting the arm 2A is not limited to those exemplified. In other words, the robot 2 has an arm 2A connected by a plurality of joints and operates by driving the joints.

[0068] The end effector 2B may include, for example, a gripping hand configured to grip the workpiece 8. The gripping hand may have a plurality of fingers. The number of fingers of the gripping hand may be two or more. The fingers of the gripping hand may have one or more joints. The end effector 2B may also include a suction hand configured to suction the workpiece 8. The end effector 2B may also include a scooping hand configured to scoop the workpiece 8. The end effector 2B may include a tool such as a drill and may be configured to perform various processes such as drilling holes in the workpiece 8. The end effector 2B is not limited to these examples and may be configured to perform various other operations. In the configuration illustrated in FIG. 10, it is assumed that the end effector 2B includes a gripping hand.

[0069] The robot 2 can control the position of the end effector 2B by operating the arm 2A. The end effector 2B may have an axis that serves as a reference for the direction in which it acts on the workpiece 8. When the end effector 2B has an axis, the robot 2 can control the direction of the axis of the end effector 2B by operating the arm 2A. The robot 2 controls the start and end of the operation in which the end effector 2B acts on the workpiece 8. The robot 2 can move or process the workpiece 8 by controlling the operation of the end effector 2B while controlling the position of the end effector 2B or the direction of the axis of the end effector 2B. In the configuration illustrated in FIG. 10, the robot 2 causes the end effector 2B to grip the workpiece 8 at the work start point 6 and moves the end effector 2B to the work target point 7. The robot 2 causes the end effector 2B to release the workpiece 8 at the work target point 7. By doing so, the robot 2 can move the workpiece 8 from the work start point 6 to the work target point 7.

[0070] <Sensor 3> As shown in FIG. 10, the robot control system 100 further includes a sensor 3. The sensor 3 detects physical information of the robot 2. The physical information of the robot 2 may include information regarding the actual position or posture of each component of the robot 2, or the speed or acceleration of each component of the robot 2. The physical information of the robot 2 may include information regarding the force acting on each component of the robot 2. The physical information of the robot 2 may include information regarding the current flowing through the motor that drives each component of the robot 2 or the torque of the motor. The physical information of the robot 2 represents the result of the actual operation of the robot 2. That is, the robot control system 100 can grasp the result of the actual operation of the robot 2 by acquiring the physical information of the robot 2.

[0071] The sensor 3 may include a force sensor or a tactile sensor that detects a force, a distributed pressure, a slip, etc. acting on the robot 2 as the physical information of the robot 2. The sensor 3 may include a motion sensor that detects the position or posture, or the speed or acceleration of the robot 2 as the physical information of the robot 2. The sensor 3 may include a current sensor that detects the current flowing through the motor that drives the robot 2 as the physical information of the robot 2. The sensor 3 may include a torque sensor that detects the torque of the motor that drives the robot 2 as the physical information of the robot 2.

[0072] The sensor 3 may be installed at a joint of the robot 2 or at a joint drive unit that drives the joint. The sensor 3 may also be installed on the arm 2A or the end effector 2B of the robot 2.

[0073] The sensor 3 outputs the detected physical information of the robot 2 to the robot control device 110. The sensor 3 detects and outputs the physical information of the robot 2 at a predetermined timing. The sensor 3 outputs the physical information of the robot 2 as time-series data.

[0074] <Camera 4> In the configuration example shown in FIG. 10, assume that the robot control system 100 includes two cameras 4. The camera 4 captures an image of an article, a human, or the like located in an influence range 5 that may affect the operation of the robot 2. The image captured by the camera 4 may include monochrome luminance information or may include luminance information of each color represented by RGB (Red, Green and Blue) or the like. The influence range 5 includes the operation range of the robot 2. Assume that the influence range 5 is a range obtained by further expanding the operation range of the robot 2 outward. The influence range 5 may be set so that the robot 2 can be stopped before a human or the like moving from the outside to the inside of the operation range of the robot 2 enters the inside of the operation range of the robot 2. The influence range 5 may be set, for example, to a range expanded by a predetermined distance from the boundary of the operation range of the robot 2 to the outside. The camera 4 may be installed so as to be able to capture an overhead view of the influence range 5, the operation range of the robot 2, or an area around these. The number of cameras 4 is not limited to two, and may be one or three or more.

[0075] <Robot control device 110> The robot control device 110 acquires the learned model 50 generated by the learned model generation device 20. Based on the image captured by the camera 4 and the learned model 50, the robot control device 110 recognizes the work object 8, the work start point 6, the work target point 7, or the like existing in the space where the robot 2 performs work. In other words, the robot control device 110 acquires the learned model 50 generated to recognize the work object 8 or the like based on the image captured by the camera 4. The robot control device 110 is also referred to as an inference device.

[0076] The robot control device 110 may be configured to include at least one processor to provide control and processing capabilities for executing various functions. Each component of the robot control device 110 may be configured to include at least one processor. A plurality of components among the components of the robot control device 110 may be realized by one processor. The entire robot control device 110 may be realized by one processor. The processor can execute programs for realizing various functions of the robot control device 110. The processor may be realized as a single integrated circuit. The integrated circuit is also referred to as an IC (Integrated Circuit). The processor may be realized as a plurality of communicably connected integrated circuits and discrete circuits. The processor may be realized based on various other known technologies.

[0077] The robot control device 110 may include a storage unit. The storage unit may include an electromagnetic storage medium such as a magnetic disk, or may include a memory such as a semiconductor memory or a magnetic memory. The storage unit stores various information and programs executed by the robot control device 110. The storage unit may be configured as a non-temporary readable medium. The storage unit may function as a work memory of the robot control device 110. At least a part of the storage unit may be configured separately from the robot control device 110.

[0078] (Operation example of the robot control system 100) The robot control device 110 (inference device) acquires a pre-trained model 50 in advance. The robot control device 110 may store the pre-trained model 50 in the storage unit. The robot control device 110 acquires an image of the work object 8 captured by the camera 4. The robot control device 110 inputs the image of the work object 8 as input information into the pre-trained model 50. The robot control device 110 acquires output information output in response to the input of the input information from the pre-trained model 50. The robot control device 110 recognizes the work object 8 based on the output information, and executes operations such as gripping or moving the work object 8.

[0079] <Parentheses> As described above, the robot control system 100 acquires the learned model 50 from the learned model generation system 1 and can recognize the work object 8 by the learned model 50.

[0080] (Use case in the component recognition model) An example of the use case of the learned model generation system 1 according to this embodiment will be described.

[0081] The actor can be the administrator of the learned model generation device 20, the user who introduces the robot 2, or the robot control device 110. The system used by the actor can be the learned model generation system 1 or the robot control system 100 that executes the pick-and-place task. The use case of each actor is exemplified below. The administrator of the learned model generation device 20 generates a general-purpose model. The user who introduces the robot 2 creates a dedicated model or registers the components to be recognized. Also, the user who introduces the robot 2 causes the robot 2 to execute the pick-and-place task. The robot control device 110 acquires the learned model 50.

[0082] <Usage pattern A> As usage pattern A, it is assumed that the user who introduces the robot 2 requests the robot control system 100 to recognize the components in a large classification such as screws or nuts when the user's own components are not included in the recognition target or even when the user's own components are included in the recognition target.

[0083] In this case, the user may use the general-purpose model without performing new learning. Therefore, the administrator of the learned model generation device 20 generates the learned model 50 as a general-purpose model. In this case, the learned model generation device 20 generates the learned model 50 using the same information as the first information used to generate the preliminary model as the second information.

[0084] <Usage pattern B> Assume a case where the robot control system 100 is requested to cause the user who introduces the robot 2 as pattern B to recognize the user's own parts.

[0085] In this case, the second information or the third information serving as learning data for recognizing the user's own parts is required. The administrator of the learned model generation device 20 generates a learned model 50 as a dedicated model by learning the second information or the third information for recognizing the user's own parts as learning data. When the learned model generation device 20 generates only the head as a dedicated model, the backbone may be generated as a general-purpose model or as a dedicated model.

[0086] (Regarding the division unit of the learned model 50) In the learned model generation system 1 described above, the learned model 50 is generated by being divided into two, i.e., the first learned model 51 and the second learned model 52. The learned model 50 is not limited to two and may be divided into three or more models. For example, when the learned model 50 has a plurality of layers, the learned model 50 may be divided into models, each having one layer. The learned model generation system 1 may generate a learned model 50 corresponding to each divided model by performing learning on a model in which each divided model is connected to a plurality of preliminary models.

[0087] For example, when the learned model 50 is divided into three models at the head, in the middle, and at the end, the learned model generation device 20 may use the middle model as a model being learned and perform learning by connecting the parts corresponding to the head and end models among the preliminary models to the middle model.

[0088] Of the learned models 50, the pre-second-learning model (head) may be divided into two or more parts. In this case, the second control unit 22 of the learned model generation device 20 may fix at least one part of the two or more parts obtained by dividing the pre-second-learning model (head). The second control unit 22 may perform learning on a model obtained by connecting the fixed part of the pre-second-learning model, a head obtained by connecting the fixed part of the pre-second-learning model and a part corresponding to the other part of the pre-second-learning model in the preliminary model, and the backbone of the preliminary model.

[0089] Of the learned models 50, the pre-first-learning model (backbone) may be divided into two or more parts. In this case, the second control unit 22 may fix at least one part of the two or more parts obtained by dividing the pre-first-learning model (backbone). The second control unit 22 may perform learning on a model obtained by connecting the fixed part of the pre-first-learning model, a backbone obtained by connecting the fixed part of the pre-first-learning model and a part corresponding to the other part of the pre-first-learning model in the preliminary model, and the head of the preliminary model.

[0090] The other part of the head and the backbone that are sequentially connected to a part of the head to be fixed, or the other part of the backbone and the head that are sequentially connected to a part of the backbone to be fixed, may be a set constructed within one preliminary model.

[0091] When fixing a part of the head, the part of the head may correspond to a lower-dimensional process than the other unfixed parts of the head. When fixing a part of the backbone, the part of the backbone may correspond to a lower-dimensional process than the other unfixed parts of the backbone. In other words, the part corresponding to the lower-dimensional process in the head or the backbone may be fixed. For example, in a model in which a CNN layer is constructed, the part corresponding to the upstream of the CNN layer may be fixed.

[0092] The learned model 50 may include a branched model as illustrated in FIG. 11. The branched model means a model in which the output from a layer branches into two or more and is input to the next layer. In the learned model 50, the first learned model 51 and the second learned model 52 may be connected in various combinations. In the branched model, the model corresponding to the first part and the model corresponding to the second part may be connected in various combinations. The branched model may be, for example, an RPN (Region Proposal Network).

[0093] (Modification of the base model according to the pre-trained model) As described above, the second control unit 22 of the learned model generation device 20 learns based on the second information in a state where the base model corresponding to the first part of the model and the pre-trained model (second pre-trained model 520) corresponding to the second part of the model are connected. In this learning, the second control unit 22 may not only change the pre-trained model corresponding to the second part of the model, but also change the base model corresponding to the first part of the model in accordance with the change of the pre-trained model. In other words, when the second control unit 22 connects the base model and the target model to execute learning, it may change various parameters in the base model set by the preliminary learning. As a result, for example, even when a base model that has undergone different learning processes is connected to a target model for learning processing, the influence of the domain gap can be reduced, and the inference accuracy of the learned model 50 can be improved. The domain gap is a phenomenon that occurs because the learning environment and the inference environment are different. That is, even for the same subject, a domain gap may occur because the acquisition environment of the image used as the learning data is different from the acquisition environment of the image used as the inference data. Therefore, in order to reduce the influence of the domain gap, it may be necessary to perform fine-tuning based on the images acquired in the inference environment (the usage environment of the robot). In other words, in order to reduce the influence of the domain gap, re-learning based on the images in the robot usage environment of the target model may be required. That is, when a part of the base model is included in the target model, when the parameters of the base model included in the target model are changed by learning based on the images in the robot usage environment, the influence of the domain gap can be reduced.

[0094] (A model in which the first part or the second part is divided into a plurality) At least one of the first part or the second part of the model may be divided into a plurality. For example, as shown in FIG. 12, each of the learned preliminary models 401 to 40N may include a first learned preliminary model 411 to 41N, a second learned preliminary model 421 to 42N, and a third learned preliminary model 431 to 43N. In the example of FIG. 12, the second learned preliminary models 421 to 42N correspond to the first part of the model. The first learned preliminary models 411 to 41N and the third learned preliminary models 431 to 43N correspond to the second part of the model.

[0095] The second control unit 22 of the learned model generation device 20 may sequentially transfer each of the second learned preliminary models 421 to 42N corresponding to the first part of the model, and learn the model connected between the first pre-learning model 510 and the third pre-learning model 530 corresponding to the second part of the model.

[0096] In the example of FIG. 12, the part connected to the input side (left side) of the second part divided into two is also referred to as an encoder. The part connected to the output side (right side) of the second part divided into two is also referred to as a decoder.

[0097] As exemplified in FIG. 12, not only the second part is divided into a plurality, but the first part may also be divided into a plurality. When the first part is divided into a plurality, the second part connected between the plurality of first parts may be learned.

[0098] (Example of operation of generating both or one of the backbone or the head) In the embodiment described in the first half of the present disclosure, in the learned model generation system 1, the head of the learned model 50 can be generated by learning a model in which the backbone of the preliminary model is connected to the learning head. On the other hand, in the learned model generation system 1, the backbone of the learned model 50 may be generated by learning a model in which the head of the preliminary model is connected to the learning backbone. Note that the learned model 50 may be used only with the head or the backbone.

[0099] When generating the backbone of the learned model 50, for example, the backbone may be generated as follows. That is, as shown in FIG. 5, the first control unit 12 of the preliminary model generation device 10 generates or acquires a plurality of pre-learning preliminary models 301 to 30N in advance. In FIG. 5, it is assumed that the configuration of the fully connected layer or the parameter size of each layer in the second pre-learning preliminary model 321 is different from the configuration of the fully connected layer or the parameter size of each layer in the second pre-learning preliminary model 32N. That is, it is assumed that the configuration of the fully connected layer or the parameter size of each layer in each of the second pre-learning preliminary models 321 to 32N is different from each other. Also, in FIG. 5, it is assumed that the configuration of the CNN layer or the filter size of each layer in the first pre-learning preliminary model 311 is the same as the configuration of the CNN layer or the filter size of each layer in the first pre-learning preliminary model 31N. That is, it is assumed that the configuration of the CNN layer or the filter size of each layer in each of the first pre-learning preliminary models 311 to 31N is the same. The configuration of the CNN layer or the filter size of each layer in each of the first pre-learning preliminary models 311 to 31N may be different from each other.

[0100] The first control unit 12 learns each pre-learning preliminary model 30 using the first information that is the same as or related to the input information input to the learned model 50 as learning data. The first control unit 12 updates each pre-learning preliminary model 30 by learning, and generates a plurality of learned preliminary models 401 to 40N. The configuration of the fully connected or CNN layer, the parameter size of each layer, or the filter size of each learned preliminary model 40 is the same as the configuration of the fully connected or CNN layer, the parameter size of each layer, or the filter size of each pre-learning preliminary model 30.

[0101] The second control unit 22 of the learned model generation device 20 acquires a learned preliminary model 40 from the preliminary model generation device 10 as a preliminary model. The second control unit 22 uses, as learning data, third information that is the same as or related to the input information input to the learned model 50, and executes learning on a model in which the head of each preliminary model is connected to a learning backbone. The head of each preliminary model corresponds to the first part of the base model. The learning backbone corresponds to the second part of the target model. The third information may be the same as or different from the first information or the second information. The third information may be configured as a set of a plurality of learning images. The second control unit 22 may perform learning on each model in which the head of each preliminary model is connected to a learning backbone, using the same set of third information as learning data, or may perform learning using different sets of third information as learning data. Further, the second control unit 22 may divide one set of third information into even smaller subsets, and perform learning using different subsets as learning data each time the head connected to the learning backbone is changed. The second control unit 22 may perform learning using the same subset as learning data when the head connected to the learning backbone is changed. That is, the second control unit 22 may execute learning on a model in which the head of each preliminary model is connected to a learning backbone based on at least one set of third information. Note that the amount of information of the third information used for learning the learned model 50 may be equal to or less than the amount of information of the first information used for learning the preliminary model. Note that the amount of information refers to, for example, the number of learning images included in the second information.

[0102] Specifically, as shown in FIG. 6, the second control unit 22 sequentially performs learning on a model obtained by transferring the heads of the respective preliminary models (the second learned preliminary models 421 to 42N) and connecting them to the first pre-learning model 510 or the first in-learning models 511 to 51(N-1). The second control unit 22 may perform learning according to the following procedure. The second control unit 22 performs learning on a model obtained by connecting the second learned preliminary model 421 and the first pre-learning model 510, and updates the first pre-learning model 510 to the first in-learning model 511. The second control unit 22 performs learning on a model obtained by connecting the second learned preliminary model 422 and the first in-learning model 511 updated in the previous step, and updates the first in-learning model 511 to the first in-learning model 512. The second control unit 22 performs learning on a model obtained by connecting the second learned preliminary model 42N and the first in-learning model 51(N-1) updated in the previous step, and updates the first in-learning model 51(N-1) to the first learned model 51N.

[0103] By executing the procedure described above, the second control unit 22 generates the first learned model 51N. The second control unit 22 applies the first learned model 51N as the first learned model 51 (the backbone of the learned model 50). The first learned model 51N generated by performing learning while connected to the first part of the base model corresponds to the second part of the target model.

[0104] As described above, in the learned model generation system 1, either or both of the backbone and the head can be generated based on learning in a model obtained by connecting the corresponding head or backbone.

[0105] Note that the relationship between the first part and the second part of a model having a backbone generated based on learning with multiple heads connected is the reverse of the relationship between the first part and the second part of a model having a head generated based on learning with multiple backbones connected in the embodiments described in the first half of the present disclosure. Specifically, in the generation of a backbone based on learning with multiple heads connected, if learning is executed by connecting the first part of each of the multiple preliminary models to the second part of the target model, it can be read as such. Conversely, also in the generation of a backbone based on learning with multiple backbones connected, as in the embodiments described in the first half of the present disclosure, if learning is executed by connecting the second part of each of the multiple preliminary models to the first part of the target model, it can also be read as such. That is, the first part and the second part of the model may be appropriately interchanged. In the learned model generation system 1, one or both of the first part and the second part of the model may be generated based on learning in a model in which the corresponding second part or first part is connected.

[0106] (Loss function) The learned model generation system 1 may set a loss function so that the output when input information is input to the generated learned model 50 approaches the output when learning data is input. In the present embodiment, cross entropy may be used as the loss function. Cross entropy is calculated as a value representing the relationship between two probability distributions. Specifically, in the present embodiment, cross entropy is calculated as a value representing the relationship between the input information and the backbone or the head.

[0107] The trained model generation system 1 learns so that the value of the loss function becomes small. In the trained model 50 generated by learning so that the value of the loss function becomes small, the output corresponding to the input of the input information can approach the output corresponding to the input of the learning data. As the loss function, for example, Discrimination Loss or Contrastive Loss may be used. Discrimination Loss is a loss function used for learning by labeling the authenticity of the generated image with a numerical value between 1 representing completely true and 0 representing completely fake.

[0108] As described above, the embodiments of the trained model generation system 1 and the robot control system 100 have been described. As embodiments of the present disclosure, in addition to a method or program for implementing a system or apparatus, an embodiment as a storage medium (for example, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a CD-RW, a magnetic tape, a hard disk, or a memory card, etc.) in which the program is recorded is also possible.

[0109] Also, the implementation form of the program is not limited to application programs such as object code compiled by a compiler and program code executed by an interpreter, and may be in the form of a program module incorporated in an operating system. Further, the program may or may not be configured such that all processing is performed only on the CPU on the control board. The program may be configured such that a part or all of it is performed by another processing unit mounted on an expansion board or an expansion unit added to the board as needed.

[0110] Embodiments according to the present disclosure have been described based on the drawings and examples. It should be noted that those skilled in the art can make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications or alterations are included in the scope of the present disclosure. For example, the functions etc. included in each component etc. can be rearranged so as not to be logically contradictory, and a plurality of components etc. can be combined into one or divided.

[0111] All of the constituent elements described in the present disclosure, and / or all of the disclosed methods, or all of the steps of the processes, can be combined in any combination except combinations in which these features are mutually exclusive. Also, each of the features described in the present disclosure can be replaced with an alternative feature that serves the same purpose, an equivalent purpose, or a similar purpose, unless explicitly negated. Therefore, unless explicitly negated, each of the disclosed features is merely an example of a comprehensive series of identical or equivalent features.

[0112] Furthermore, the embodiments according to the present disclosure are not limited to any specific configuration of the above-described embodiments. The embodiments according to the present disclosure can be extended to all of the novel features described in the present disclosure, or combinations thereof, or all of the novel methods, or process steps, or combinations thereof.

[0113] In the present disclosure, descriptions such as "first" and "second" are identifiers for distinguishing the relevant configuration. The configurations distinguished by the descriptions such as "first" and "second" in the present disclosure can have their numbers in the relevant configuration exchanged. For example, the first information can have the "first" and "second" which are identifiers exchanged with the second information. The exchange of the identifiers is performed simultaneously. The relevant configuration is still distinguishable after the exchange of the identifiers. The identifiers can be deleted. The configuration with the identifiers deleted is distinguished by reference signs. Based only on the descriptions of the identifiers such as "first" and "second" in the present disclosure, the order of the relevant configuration should not be interpreted, nor should it be used as the basis for the existence of an identifier with a smaller number.

Description of Reference Signs

[0114] 1 Learned model generation system 10 Preliminary model generation device (12: Control unit) 20 Learned model generation device (22: Control unit) 30(301~30N) Pre-learning preliminary models (31(311~31N): First pre-learning preliminary model, 32(321~32N): Second pre-learning preliminary model) 40(401~40N) Learned preliminary models (41(411~41N): First learned preliminary model, 42(421~42N): Second learned preliminary model, 431~43N: Third learned preliminary model) 50 Learned models (51: First learned model, 52: Second learned model, 510: First pre-learning model, 520: Second pre-learning model) 100 Robot control system (2: Robot, 2A: Arm, 2B: End effector, 3: Sensor, 4: Camera, 5: Influence range of the robot, 6: Work start platform, 7: Work target platform, 8: Work object, 110: Robot control device (inference device)

Claims

1. A method for generating a learned model, comprising: a processor generating a learned model that outputs a recognition result of a recognition target included in input information based on a plurality of models each having at least one of a first part and a second part. In the generation of the learned model, the processor obtains a plurality of base models each having a part corresponding to the first part, which are learned based on at least one set of first information related to the input information. the processor obtains a target model each having a part corresponding to the second part, which is learned based on at least one set of second information related to the input information while being connected to the plurality of base models. the processor generates a learned model having at least the target model having the part corresponding to the second part. A method for generating a learned model.

2. A method for generating a learned model, comprising: a processor generating a learned model that outputs a recognition result of a recognition target included in input information based on a plurality of models each having at least one of a first part and a second part. In the generation of the learned model, the processor obtains a plurality of base models each having a part corresponding to the first part, which are learned based on at least one set of first information related to the input information. the processor obtains a target model each having a part corresponding to the second part, which is learned based on at least one set of second information related to the input information while being connected to the plurality of base models. the processor generates a learned model having at least the target model having the part corresponding to the second part, and at least one of the plurality of base models has a model configuration different from that of other base models. A method for generating a learned model.

3. The method for generating a learned model according to claim 1 or 2, wherein the second information is the same as the first information.

4. The method for generating a learned model according to claim 1 or 2, wherein the second information is different from the first information.

5. The method for generating a learned model according to claim 1 or 2, further comprising generating, as the learned model, a model in which a target model corresponding to the first part learned based on third information related to the input information is connected to a target model corresponding to the second part.

6. The method for generating a learned model according to claim 5, wherein the third information is the same as the first information.

7. The method for generating a learned model according to claim 6, wherein the third information is different from the first information.

8. The first part of the model corresponds to a backbone for extracting feature amounts of the recognition target, The method for generating a learned model according to claim 1 or 2, wherein the second part of the model corresponds to a head for outputting the recognition result based on the extraction result of the feature amounts.

9. The first part of the model corresponds to a head for outputting the recognition result based on the extraction result of the feature amounts of the recognition target, The method for generating a learned model according to claim 1 or 2, wherein the second part of the model corresponds to a backbone for extracting the feature amounts.

10. The method for generating a learned model according to claim 1 or 2, wherein the learned model includes a branch model.

11. The method for generating a learned model according to claim 1 or 2, further comprising: the processor learns based on the second information in a state where a base model corresponding to the first part of the model and a pre-learning model corresponding to the second part of the model are connected, thereby changing the base model to match the pre-learning model and generating a model obtained by changing the pre-learning model as the target model.

12. In the generation of the learned model, the method for generating a learned model according to claim 1 or 2, wherein the processor generates the plurality of base models by learning using the same set of the first information.

13. Each of the plurality of models further has a third part. In the generation of the learned model, the processor obtains a target model having parts corresponding to the second part and the third part respectively, which are learned based on at least one set of second information related to the input information in a state where the target model is connected to the plurality of base models, and the processor generates a learned model having a target model having at least parts corresponding to the second part and the third part. The method for generating a learned model according to claim 1 or 2.

14. A control unit is provided for generating a learned model that outputs a recognition result of a recognition target included in input information based on a plurality of models each having at least one of a first part and a second part. In the generation of the learned model, the control unit obtains a plurality of base models each having a part corresponding to the first part, which are learned based on at least one set of first information related to the input information, and obtains a target model having a part corresponding to the second part, which is learned based on at least one set of second information related to the input information in a state where the target model is connected to the plurality of base models, and generates a learned model having a target model having at least a part corresponding to the second part. A learned model generation device.

15. A control unit that generates a learned model that outputs a recognition result of a recognition target included in input information based on a plurality of models each having at least one of a first part and a second part. In generating the learned model, the control unit acquires a plurality of base models each having a part corresponding to the first part, which are learned based on at least one set of first information related to the input information, and acquires a target model having a part corresponding to the second part, which is learned based on at least one set of second information related to the input information while being connected to the plurality of base models, and generates a learned model having the target model having at least a part corresponding to the second part. At least one of the plurality of base models has a model configuration different from that of other base models. Learned model generation device.

Citation Information

Patent Citations

  • Tire damage detection and recognition method and device

    CN112613375A

  • Weak supervision semantic segmentation model training method and device, storage medium and terminal

    CN113159049A

  • Method for monitoring and diagnosing periodic vibrating phenomenon

    JP1999183246A

  • Biosensing method and device, system, electronic device, and storage medium

    JP2020522764A

  • Control method of robot system, manufacturing method of articles, control program, recording medium, and robot system

    JP2021013996A