Information Processing Apparatus and Information Processing Method

The CANN model addresses the challenges of domain offset and label distribution differences in transfer learning by using collaborative training and self-supervised clustering, resulting in improved detection accuracy and reduced negative transfer.

JP7690879B2Active Publication Date: 2025-06-11FUJITSU LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021204773
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-20
Filing Date
2021-12-17
Publication Date
2025-06-11
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

Existing transfer learning methods face challenges in accurately classifying unlabeled samples in a target domain when there is a domain offset and differences in label distributions between the source and target domains.

Method used

An information processing apparatus and method that utilize a Coupled Approximation Neural Network (CANN) model, which includes a first determination unit and a second determination unit for collaborative training. The CANN model extracts features using self-supervised clustering and adjusts for domain differences by exchanging information between the units.

Benefits of technology

The CANN model significantly improves detection accuracy in transfer learning scenarios by effectively aligning shared classes between domains and reducing negative transfer, while preserving the inherent features of the samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690879000016
    Figure 0007690879000016
  • Figure 0007690879000017
    Figure 0007690879000017
  • Figure 0007690879000018
    Figure 0007690879000018
Patent Text Reader

Abstract

To provide an information processing device and an information processing method.SOLUTION: An information processing device classifies a second sample without a label of a second domain different from a first domain based on a first sample with a label of the first domain, wherein a class of the second sample is a subset or a whole set of a class for the first sample. The information processing device includes: a first determination section for determining whether or not the class of the first sample is a common class for the first domain and the second domain; and a second determination section for determining the common class to which the second sample belongs. The first determination section and the second determination section execute a joint training through exchanging information. According to an information processing technology of the present disclosure, a learning effect for transfer learning can be significantly enhanced.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of information processing. In particular, embodiments of the present disclosure relate to an information processing apparatus, an information processing method, and a computer-readable storage medium for transfer learning between different domains.

Background Art

[0002] In recent years, deep learning has made great progress in many machine learning tasks and applications. As the application scenarios of machine learning increase, the performance of supervised machine learning is improving, but a large number of labeled samples are required for training. However, labeling samples, that is, attaching labels to samples, is time-consuming and laborious. Therefore, transfer learning, which can apply knowledge learned in one domain (hereinafter referred to as the source domain) to another different and related domain (hereinafter referred to as the target domain), has been attracting increasing attention.

[0003] Also, many machine learning tasks are based on the assumption that the training sample set (that is, a set of labeled samples) and the test set (that is, a set of unlabeled samples) have the same sample distribution. However, in actual applications, especially in transfer learning, this assumption may not hold. In transfer learning, there is a strong domain offset in samples collected from different domains. Factors of the domain offset include, for example, changes in the environment, changes in the sampling method, and different data patterns of samples. Also, different sample label distributions (that is, class distributions of samples) between the source domain and the target domain also affect the training of the model. The above factors affect the accuracy of the final transfer learning model.

Summary of the Invention

Problems to be Solved by the Invention

[0004] The following provides a brief overview of the present disclosure to facilitate a basic understanding of the aspects thereof. Note that this brief overview is not an exhaustive overview of the present disclosure, nor is it intended to specifically identify the key points or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Instead, it aims to simply explain the concepts in a simple form as a prelude to the more detailed description to follow.

[0005] In view of the above-described problems of the prior art, an object of the present disclosure is to provide an information processing technology for transfer learning between different domains that can accurately and quickly detect samples without labels.

Means for Solving the Problem

[0006] To achieve the object of the present disclosure, in one aspect of the present disclosure, there is provided an information processing apparatus for classifying a second sample having no label of a second domain different from the first domain based on a first sample having a label of the first domain, wherein the class of the second sample is a subset or the entire set of the classes of the first sample, the information processing apparatus including: a first determination unit that determines whether the class of the first sample is a shared class between the first domain and the second domain; and a second determination unit that determines the shared class to which the second sample belongs, and the first determination unit and the second determination unit execute collaborative training by exchanging information.

[0007] In another aspect of the present disclosure, there is provided an information processing method for classifying a second sample having no label of a second domain different from the first domain based on a first sample having a label of the first domain, wherein the class of the second sample is a subset or the entire set of the class of the first sample. The information processing method includes: a first determination unit determining whether the class of the first sample is a shared class between the first domain and the second domain; a second determination unit determining the shared class to which the second sample belongs; and performing collaborative training on the first determination unit and the second determination unit by exchanging information.

[0008] In yet another aspect of the present disclosure, there is further provided a computer program capable of realizing the above information processing method.

[0009] Further provided is a computer program product in the form of at least a computer-readable medium having recorded thereon computer program code for realizing the above information processing method. Specifically, in yet another aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a computer, it is an information processing method for classifying a second sample having no label of a second domain different from the first domain based on a first sample having a label of the first domain, wherein the class of the second sample is a subset or the entire set of the class of the first sample. The information processing method includes: a first determination unit determining whether the class of the first sample is a shared class between the first domain and the second domain; a second determination unit determining the shared class to which the second sample belongs; and performing collaborative training on the first determination unit and the second determination unit by exchanging information. Further provided is a storage medium for realizing the information processing method.

[0010] According to the information processing technology related to the present disclosure, an information processing technology based on a new transfer learning model is provided, which can significantly improve the detection accuracy compared with the conventional transfer learning model.

Brief Description of the Drawings

[0011] In order to more easily understand the above and other objects, features and advantages of the present disclosure, the embodiments of the present disclosure will be described below with reference to the drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0012] The exemplary embodiments of the present disclosure will be described below with reference to the drawings. When indicating the elements of the drawings using reference numerals, the same elements in different drawings are denoted by the same reference numerals. Also, in the following description of the present disclosure, detailed descriptions of known functions and configurations are omitted in the present disclosure.

[0013] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. Unless the context indicates otherwise, the singular forms used in this specification also include the plural forms. In addition, the terms "comprising," "including," and "having" as used in this specification mean the presence of the described features, entities, operations, and / or components, but do not exclude the presence or addition of one or more other features, entities, operations, and / or components.

[0014] Unless otherwise defined, all terms, including technical and scientific terms used in this specification, shall have the same meaning as commonly understood by one of ordinary skill in the art. Also, terms defined as in a commonly used dictionary shall be interpreted to have a meaning that coincides with their meaning in the context of the relevant art, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.

[0015] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. The present disclosure can be practiced without some or all of these specific details. In other instances, only components closely related to the configuration according to the present disclosure are illustrated in order to avoid obscuring the present disclosure with unnecessary details, and other details not related to the present disclosure are omitted.

[0016] The problem that the present disclosure aims to solve is the problem of partial domain adaptation (PDA) of transfer learning. In the PDA problem, a source domain (hereinafter also referred to as the first domain) contains labeled samples (hereinafter also referred to as the first samples), and a target domain (hereinafter also referred to as the second domain) contains unlabeled samples (hereinafter also referred to as the second samples). Also, in the PDA problem, it is defined that the classes of the second samples are a subset or the entire set of the classes of the first samples. In other words, it is defined that all the classes of the second samples in the target domain are included within the classes of the first samples in the source domain. Also, in this specification, the classes of the second samples in the target domain that are included within the classes of the first samples in the source domain are referred to as the shared classes of the source domain, and the remaining classes among the classes of the first samples in the source domain are referred to as the non-shared classes. As described above, the number of non-shared classes is a natural number.

[0017] The PDA problem is a problem of classifying samples (second samples) without labels in a target domain (second domain) based on samples (first samples) with labels in a source domain (first domain). Specifically, for the PDA problem, a sample set D s ={(x i s , y i s )} (the label of sample x i , that is, the class is y i ) contains m s samples and n s classes, the marginal distribution of sample x i is ρ s , and a sample set D t ={(x i t )} of the target domain contains m t samples and n t classes, and the marginal distribution of sample x t is ρt Assume that it is so. Also, as described above, n t ≦n s , that is, the classes of the source domain can completely cover the classes of the target domain. Also, the closure degree of the source domain is ζ = n t / n s is defined as. The smaller the closure degree ζ, the fewer the shared classes, and the higher the risk of negative transfer.

[0018] In the PDA problem, the main issue is the alignment of the shared classes between the source domain and the target domain. Obviously, as the closure degree ζ of the source domain decreases, the probability that the samples in the target domain match the non-shared classes increases, leading to negative transfer. In the prior art, many methods have been proposed to solve the PDA problem. These methods mainly learn the domain-invariant features between the source domain and the target domain (for example, through domain adversarial training) to optimize the classification loss in the source domain in order to realize the model training of the target domain. However, in the PDA problem, since the non-shared classes are included in the source domain, the directly extracted domain-invariant features are still not essentially aligned between the two domains (that is, the source domain and the target domain). Therefore, as the closure degree of the source domain decreases, the domain-invariant features become increasingly inaccurate, leading to negative transfer. On the other hand, such domain-invariant features may destroy the inherent features of the original samples, such as the high similarity of similar samples and the constraints of the potential conditional distribution.

[0019] Therefore, one objective of the information processing technology according to the embodiments of the present disclosure is to solve the PDA problem, that is, to train the general model for classifying the samples of the target domain without knowing the closure degree of the source domain so that the general model can exhibit excellent detection performance with a wide range of closure degrees.

[0020] In the present disclosure, a new model for solving the PDA problem is provided, which may be referred to as a Coupled Approximation Neural Network (CANN) model. Different from the conventional methods, the CANN model according to the present disclosure focuses on relaxing the domain-invariant constraints widely used in the prior art when solving the PDA problem and extracting and utilizing the original features of the samples.

[0021] The following will describe the information processing apparatus 100 according to the embodiments of the present disclosure in detail with reference to FIGS. 1 to 5.

[0022] FIG. 1 is a block diagram showing the configuration of the information processing apparatus 100 according to the embodiments of the present disclosure. FIG. 2 is a schematic diagram showing the principle of the information processing apparatus 100 according to the embodiments of the present disclosure. FIG. 3 is a schematic diagram showing the configuration of the information processing apparatus 100 according to the embodiments of the present disclosure.

[0023] The information processing apparatus 100 according to the embodiments of the present disclosure classifies a second sample having no label of a second domain different from the first domain based on a first sample having a label of the first domain. Here, the class of the second sample is a subset or the entire set of the class of the first sample. The information processing apparatus 100 may include a first determination unit 101 and a second determination unit 102. The first determination unit 101 determines whether the class of the first sample is a shared class between the first domain and the second domain. The second determination unit 102 determines the shared class to which the second sample belongs. Here, the first determination unit 101 and the second determination unit 102 perform joint training by exchanging information.

[0024] The information processing apparatus 100 according to the embodiments of the present disclosure may be realized by using the above CANN model.

[0025] Specifically, the CANN model according to the present disclosure extracts potential features of samples in both the source domain (i.e., the first domain) and the target domain (i.e., the second domain) based on the self-supervised clustering algorithm. The CANN model according to the present disclosure may include two parallel sub-network models, namely, the Target to Source Approximation Neural Network (T-S ANN) model and the Target to Target Approximation Neural Network (T-T ANN) model.

[0026] In an embodiment of the present disclosure, the first determination unit 101 may be implemented using the above T-S ANN model, and the second determination unit 102 may be implemented using the above T-T ANN model.

[0027] In an embodiment of the present disclosure, the first determination unit 101 and the second determination unit 102 may be implemented by a convolutional neural network model (CNN). Also, in an embodiment of the present disclosure, the first determination unit 101 and the second determination unit 102 may be implemented by a CNN model having the same structure.

[0028] In other words, both the T-S ANN model for implementing the first determination unit 101 and the T-T ANN model for implementing the second determination unit 102 may be implemented by a CNN model and may have the same model structure.

[0029] Also, in the embodiments of the present disclosure, the first sample and the second sample may be images. That is, the sample (the first sample) in the source domain is a labeled image, and the sample (the second sample) in the target domain is an unlabeled image. For example, as shown in FIGS. 3 and 4, the source domain may be a domain of comic images, and the target domain may be a domain of photographic images. Also, as shown in FIGS. 3 and 4, for example, the images in the source domain may have four types of labels, i.e., four classes of bed, chair, bag, and TV, and the images in the target domain may have two types of labels, i.e., two classes of bed and chair.

[0030] As described above, in the problem of PDA, the specific classes of samples in the target domain, such as photographic images, are unknown, but the classes are necessarily included in the classes of the source domain. The information processing apparatus 100 according to the embodiments of the present disclosure can determine the classes of samples in the target domain based on the knowledge regarding the samples in the source domain.

[0031] As a non-limiting example, each of the first determination unit 101 and the second determination unit 102 may use a ResNet-50 model pre-trained based on the same-structured ImageNet image database as a backbone network.

[0032] Since the convolutional neural network (CNN) model is known to those skilled in the art, for the sake of brevity, the details thereof are omitted in this specification.

[0033] Note that although the exemplary embodiments of the present disclosure are described in this specification with images as examples of samples, the scope of the present disclosure is not limited thereto. Based on the teachings of the present disclosure, the information processing apparatus according to the present disclosure may be applied to other machine learning application fields such as voice analysis, natural language processing, weather forecasting, etc. in addition to image processing.

[0034] As shown in FIG. 2, the triangular mark represents a sample in the source domain, and the circular mark represents a sample in the target domain. As shown in FIG. 2, the T-S ANN model used as the first determination unit 101 may find a source domain adjacent sample corresponding to a target domain sample based on feature similarity and realize clustering from the target domain to the source domain using a self-teaching algorithm. Based on the assumption that samples of the same class between the two domains have high similarity, the T-S ANN model can more accurately evaluate the class difference between the two domains.

[0035] Also, as shown in FIG. 2, the T-T ANN model used as the second determination unit 102 may find an adjacent sample set from a specific target domain sample to other target domain samples based on feature similarity in order to complete clustering. The T-T ANN model can extract the potential features of the target domain samples to the maximum extent and evaluate the reliability (uncertainty) of the target domain samples.

[0036] Also, the following will explain ν and ω shown in FIG. 2 in detail.

[0037] In the information processing apparatus 100 realized by the CANN model according to the present disclosure, the first determination unit 101 and the second determination unit 102 may be realized as a T-S ANN model and a T-T ANN model which are parallel sub-networks.

[0038] In the embodiment of the present disclosure, in the joint training, the first determination unit 101 supplies the first information regarding the shared class (represented by ν shown in FIGS. 2 to 4) to the second determination unit 102, and the second determination unit 102 supplies the second information regarding the second sample (represented by ω shown in FIGS. 2 to 4) to the first determination unit 101.

[0039] In an embodiment of the present disclosure, the first determination unit 101 and the second determination unit 102 are not independent of each other and perform joint training in a combined manner. In the forward propagation of the joint training, the T-S ANN model used as the first determination unit 101 may evaluate the class difference between the two domains, and the T-T ANN model used as the second determination unit 102 may evaluate the reliability of the target domain samples. Before the backpropagation of the joint training, these two sub-networks may simultaneously improve the performance of these two sub-networks by exchanging information regarding the class difference between the domains (i.e., the first information ν) and information regarding the reliability of the target domain samples (i.e., the second information ω) with each other. According to the information regarding the class difference between the domains (i.e., the first information ν) received from the T-S ANN model, the T-T ANN model can weaken the influence of negative transfer by non-shared classes. According to the information regarding the reliability of the target domain samples (i.e., the second information ω) received from the T-T ANN model, the T-S ANN model can filter the target domain samples and use more reliable target domain samples to evaluate the class difference between the two domains.

[0040] As shown in FIG. 3, G ts-f and G tt-f respectively represent the features of the samples extracted by the feature extraction layer of the T-S ANN model used as the first determination unit 101 and the T-T ANN model used as the second determination unit 102. Also, G ts-y and G tt-y respectively represent the classification outputs of the T-S ANN model used as the first determination unit 101 and the T-T ANN model used as the second determination unit 102.

[0041] As shown in FIG. 3, L cts represents the self-supervised clustering loss from the target domain samples to the source domain samples and is used to move the target domain samples to the corresponding source domain adjacent samples. L cttrepresents the self-supervised clustering loss from the target domain sample to the target domain sample, and is used to move the target domain sample to the corresponding target domain adjacent sample. L cc represents the error loss in the T-T ANN model used as the second determination unit 102, and is used to perform error correction on the samples classified as non-shared classes in the target domain so as to eliminate the negative transfer of non-shared classes.

[0042] As shown in FIG. 3, ν may represent the first information regarding the shared class that the T-S ANN model used as the first determination unit 101 supplies to the T-T ANN model used as the second determination unit 102. In an embodiment of the present disclosure, the first information ν may represent the probability that the class of the first sample (for example, an image in the source domain) is a shared class. As shown in FIG. 3, the probability that the class of the first sample in the source domain is a shared class is shown in the form of a histogram, and the probability that the class of the bed and the class of the chair are shared classes is the highest.

[0043] In an embodiment of the present disclosure, L cc may be expressed as the entropy loss combined with ν.

[0044] Also, as shown in FIG. 3, ω represents the second information regarding the second sample that the T-T ANN model used as the second determination unit 102 supplies to the T-S ANN model used as the first determination unit 101. In an embodiment of the present disclosure, the second information ω may represent the reliability of the second sample. As shown in FIG. 3, the second samples are shown in descending order of reliability.

[0045] In an embodiment of the present disclosure, L cts may be expressed as the movement loss combined with ω.

[0046] In an embodiment of the present disclosure, the first determination unit 101 may adjust the second sample in the collaborative training based on the second information ω received from the second determination unit 102. Also, in an embodiment of the present disclosure, the second determination unit 102 may adjust the shared class in the collaborative training based on the first information ν received from the first determination unit 101.

[0047] FIG. 4 is a schematic diagram showing collaborative training of an information processing apparatus according to an embodiment of the present disclosure.

[0048] As shown in FIG. 4, in collaborative training, in the T-S ANN model used as the first determination unit 101, F s may represent the features of the samples of the source domain that have been extracted, (Outer 1) becomes TIFF0007690879000001.tif14127, and φ ts may represent the similarity between the features of the samples of the target domain that have been extracted and F s and, (Outer 2) TIFF0007690879000002.tif16127 becomes.

[0049] Also, for the samples x i (1 ≤ i ≤ m t ) of the target domain, the similarity ρ j (1 ≤ j ≤ m s ) with the samples x i,j ts of the source domain may be calculated according to the following formula (1).

Equation

[0050] In formula (1), ||.|| 2 represents the 2-norm, and T represents the transpose.

[0051] According to the above formula (1), for the sample x of the target domain iRegarding the similarity φ with samples in the source domain i ts may be represented by the following formula (2).

Equation

[0052] Therefore, the loss function L representing the self - supervised clustering loss from the target - domain sample to the source - domain sample cts may be calculated according to the following formula (3).

Equation

[0053] In formula (3), D t and D s represent the sample set of the target domain and the sample set of the source domain respectively, and ω i is the second information provided by the T - T ANN model used as the second determination unit 102. The following will explain this information in detail.

[0054] In the embodiment of the present disclosure, the T - S ANN model used as the first determination unit 101 may estimate the class difference between the two domains by taking the statistics of the label distribution (i.e., class distribution) of the adjacent samples in the target domain. Specifically, for each class l (1 ≤ l ≤ n s ) in the source domain, the weight (probability) ν l which is the shared class may be calculated according to the following formula (4).

Equation

[0055] In formula (4), μ l represents the statistical result regarding class l in the source domain, (Outer 3) TIFF0007690879000007.tif17127 becomes. Sample x of the target domain i For, the cumulative weight μ l may be updated according to the following formula (5).

Equation

[0056] In Equation (5), σ = argmaxφ i ts becomes.

[0057] In the embodiments of the present disclosure, when the scanning of all target domain samples is completed, the first information ν is reset.

[0058] Also, as shown in FIG. 4, in the collaborative training, in the T-T ANN model used as the second determination unit 102, F t may represent the features of the extracted samples of the target domain, (External 4) TIFF0007690879000009.tif14127 becomes, and φ tt may represent the similarity between the features of the extracted target domain samples and F t . Here, when calculating the adjacent samples of the target domain, it is necessary to consider excluding the samples of the target domain itself. Therefore, (External 5) TIFF0007690879000010.tif17127 becomes.

[0059] Also, for the sample x of the target domain i (1 ≤ i ≤ m t ), for other samples x of the target domain j (1 ≤ j ≤ m t , j ≠ i), the similarity ρ i,j tt may be calculated according to the following formula (6).

Equation

[0060] Therefore, for the sample x of the target domain i the similarity φ with other samples in the target domain i tt may be represented by the following formula (7).

Equation

[0061] Therefore, the loss function L representing the self-supervised clustering loss from a target domain sample to a target domain sample ctt may be calculated according to the following formula (8).

Equation

[0062] When estimating the reliability of the sample x of the target domain i the more difficult it is to classify the sample of the target domain or the lower its reliability, the easier it is for the sample to be isolated in the feature space. Therefore, in the embodiments of the present disclosure, the T-T ANN model used as the second determination unit 102 i for the sample x of the target domain calculates the uniformly distributed Kullback-Leibler divergence (i.e., relative entropy) of its similarity matrix to estimate the reliability of the sample x of the target domain i This may be represented by the second piece of information ω i and the second piece of information ω i may be calculated according to the following formula (9).

Equation

[0063] In formula (9), θ may be a minimum value to avoid log0.

[0064] Thus, the error loss L in the T-T ANN model used as the second determination unit 102 cc may be calculated according to the following formula (10).

Equation

[0065] In Equation (10), y i,j cor = y i,j * ν j Here, y i,j represents the probability that the label (class) of the sample x of the target domain output by the T-T ANN model used as the second determination unit 102 is j. i

[0066] In the embodiments of the present disclosure, considering the constraints of the hardware computing power, the training of the CANN model (including the T-S ANN model used as the first determination unit 101 and the T-T ANN model used as the second determination unit 102) for realizing the information processing apparatus 100 may be performed batch by batch. As shown in FIG. 4, it is assumed that the number of samples of the target domain in each batch is m. For each batch, the above formulas (1) to (10) are similarly applicable, and it is only necessary to replace mt with m.

[0067] Using the loss functions L cts , L ctt and L cc given in the above formulas (3), (8), and (10), joint training can be performed on the T-S ANN model and the T-T ANN model included in the CANN model to obtain the final information processing apparatus 100.

[0068] ​As shown in FIG. 4, in the joint training, the T-S ANN model used as the first determination unit 101 supplies the first information ν regarding the shared class to the T-T ANN model used as the second determination unit 102, and the T-T ANN model used as the second determination unit 102 supplies the second information ω regarding the samples of the target domain to the T-S ANN model used as the first determination unit 101.

[0069] As described above, the higher the reliability of the samples in the target domain, the easier it is to correctly match the shared class. Conversely, the lower the reliability of the samples in the target domain, the more difficult it is to correctly match the shared class, that is, the higher the uncertainty. In an embodiment of the present disclosure, the second information ω may reflect the reliability (uncertainty) of the samples in the target domain. In this regard, as shown in FIG. 4, in the joint training, the T-S ANN model used as the first determination unit 101 may use the second information ω to increase the reliability gap between high-reliability samples and low-reliability samples and control the concentration degree of the sample distribution. In other words, according to the second information ω, the distribution of reliable target domain samples can be made more concentrated (sharp), while the distribution of unreliable target domain samples can be made more dispersed (loose), and a more reliable result can be obtained.

[0070] Also, in order to reduce the influence of negative transfer in the target domain of the non-shared class, in the joint training, in an embodiment of the present disclosure, the T-T ANN model used as the second determination unit 102 uses the first information ν from the T-S ANN model used as the first determination unit 101. In an embodiment of the present disclosure, the first information ν may reflect the probability (weight) that a specific class in the source domain is the shared class. In this regard, as shown in FIG. 4, in the joint training, the T-T ANN model used as the second determination unit 102 may use the first information ν to increase the probability (entropy loss) of the shared class and decrease the probability (entropy loss) of the non-shared class in order to weaken the influence of negative transfer in the target domain of the non-shared class.

[0071] In addition, the first information ν and the second information ω that are exchanged with each other in the joint training process of the T-S ANN model used as the first determination unit 101 and the T-T ANN model used as the second determination unit 102 can improve the performance of the T-S ANN model and the T-T ANN model simultaneously.

[0072] The information processing apparatus 100 according to the embodiment of the present disclosure is realized by a CANN model, and can significantly improve the accuracy when solving the PDA problem as compared with the methods of the prior art.

[0073] For example, the experiment is conducted based on the Office-Home image database that is commonly used in the prior art and includes four domains: art image Art(Ar), clipart image Clipart(Cl), product image Product(Pr), and real-world photo image Real World(Rw). Each domain contains 65 classes. For each domain, the top 25 classes in alphabetical order are used as shared classes, and the remaining classes are used as non-shared classes. Thereby, the information processing apparatus 100 according to the embodiment of the present disclosure is trained.

[0074] As can be seen from the experimental results, the average detection accuracy regarding the samples of the target domain of the information processing apparatus 100 according to the embodiment of the present disclosure is at least 4% higher than that of the methods of the prior art and can reach 75.5%.

[0075] FIG. 5 is a schematic diagram showing the training effect of the information processing apparatus according to the embodiment of the present disclosure.

[0076] As shown in FIG. 5(a), the three curves from top to bottom are L cc , L ctt and L ctsshows the loss, and as the iterations of the joint training progress, the loss tends to stabilize. Also, as shown in (b) of FIG. 5, it shows the weights of the 65 classes in the domain that are shared classes respectively. According to the joint training, the shared classes (for example, the top 25 classes) and non-shared classes in the source domain can be determined more accurately.

[0077] Accordingly, the present disclosure further provides an information processing method for transfer learning between different domains.

[0078] FIG. 6 is a flowchart showing an information processing method 600 according to an embodiment of the present disclosure.

[0079] The information processing method 600 starts from step S601. Next, in step S602, the first determination unit determines whether the class of the first sample is a shared class between the first domain and the second domain. In the embodiment of the present disclosure, the processing in step S602 may be realized by the first determination unit 101 described with reference to FIGS. 1 to 5 above, and the detailed description thereof is omitted here.

[0080] Next, in the determination step S603, the second determination unit determines the shared class to which the second sample belongs. In the embodiment of the present disclosure, the processing in step S302 may be realized by the second determination unit 102 described with reference to FIGS. 1 to 5 above, and the detailed description thereof is omitted here.

[0081] Next, in step S604, joint training is performed on the first determination unit and the second determination unit by exchanging information. In the embodiment of the present disclosure, the processing in step S604 may be realized by the interaction between the first determination unit 101 and the second determination unit 102 described with reference to FIGS. 1 to 5 above, and the detailed description thereof is omitted here.

[0082] Finally, the information processing method 600 ends at step S605.

[0083] According to the information processing technology of the present disclosure, the detection accuracy of transfer learning can be significantly improved without impairing the unique features of the original sample.

[0084] FIG. 7 is a schematic diagram showing the configuration of a general-purpose device for realizing the information processing method according to an embodiment of the present disclosure. The general-purpose device 700 may be, for example, a computer system. Note that the general-purpose device 700 is merely an example and does not limit the usage range or functions of the information processing method and the information processing apparatus of the present disclosure. Further, the general-purpose device 700 does not depend on the constituent elements or combinations thereof in the above information processing method and information processing apparatus.

[0085] In FIG. 7, a central processing unit (CPU) 701 executes various processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 stores data necessary for the CPU 701 to execute various processes as needed. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output interface 705 is also connected to the bus 704.

[0086] The input unit 706 (including a keyboard, a mouse, etc.), the output unit 707 (including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.), the storage unit 708 (including, for example, a hard disk, etc.), and the communication unit 709 (including a network interface card such as a LAN card, a modem, etc.) are connected to the input / output interface 705. The communication unit 709 executes communication processing via a network such as the Internet. If necessary, the driver 710 may be connected to the input / output interface 705. The removable medium 711 is, for example, a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., and is set up in the driver 710 as necessary, and the computer program read from it is installed in the storage unit 708 as necessary.

[0087] When the above processing is implemented by software, a program constituting the software is installed via a network such as the Internet or a storage medium such as the removable medium 711.

[0088] Note that these storage media are not limited to the removable medium 711 shown in FIG. 7 that stores a program and provides the program to the user separately from the device. The removable medium 711 includes, for example, a magnetic disk (including a floppy disk), an optical disk (including a compact disc read-only memory (CD-ROM) and a digital versatile disc (DVD)), a magneto-optical disk (mini disc (MD) (registered trademark)), and a semiconductor memory. Alternatively, the storage medium may be a ROM 702, a hard disk included in the storage unit 708, etc., which stores a program and is provided to the user together with the device including them.

[0089] Furthermore, the present disclosure further provides a program product storing machine-readable instruction codes. When the instruction codes are read and executed by a device, the information processing method according to the present disclosure described above can be executed. Therefore, the various storage media described above storing this program product are also included in the scope of the present disclosure.

[0090] The above has described specific embodiments of the apparatus and / or method of the embodiments of the present disclosure by elaborating block diagrams, flowcharts, and / or embodiments in detail. If one or more functions and / or operations are included in these block diagrams, flowcharts, and / or embodiments, each function and / or operation in these block diagrams, flowcharts, and / or embodiments may be implemented individually and / or collectively by hardware, software, firmware, or any combination thereof. In one embodiment, some parts of the subject matter described in this specification may be implemented by application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), or other integrated forms. Note that all or some aspects of the embodiments described in this specification may be equivalently implemented in the form of one or more computer programs executed by one or more computers in an integrated circuit (for example, in the form of one or more computer programs executed by one or more computer systems), in the form of one or more programs executed by one or more processors (in the form of one or more programs executed by one or more microprocessors), in the form of firmware, or in substantially any combination of these forms. Also, according to the content disclosed in this specification, all the circuits for designing the present disclosure and / or the codes for editing the software and / or firmware of the present disclosure are within the capabilities of those skilled in the art.

[0091] Regarding the embodiments including the above-described respective examples, the following appendices are further disclosed, but are not limited to these appendices. (Appendix 1) An information processing apparatus that classifies a second sample having no label of a second domain different from the first domain based on a first sample having a label of the first domain, wherein a class of the second sample is a subset or the entire set of the class of the first sample. A first determination unit that determines whether a class of the first sample is a shared class between the first domain and the second domain. A second determination unit that determines a shared class to which the second sample belongs, and includes: The first determination unit and the second determination unit execute joint training by exchanging information. (Appendix 2) The information processing apparatus according to Appendix 1, wherein the first determination unit and the second determination unit are realized by a convolutional neural network model. (Appendix 3) The information processing apparatus according to Appendix 2, wherein the first determination unit and the second determination unit are realized by a convolutional neural network model having the same structure. (Appendix 4) The information processing apparatus according to Appendix 1, wherein the first sample and the second sample are images. (Appendix 5) In the joint training, the first determination unit supplies first information regarding the shared class to the second determination unit, and the second determination unit supplies second information regarding the second sample to the first determination unit. (Appendix 6) The information processing apparatus according to Appendix 5, wherein the first information indicates a probability that the class of the first sample is the shared class. (Appendix 7) The information processing apparatus according to Appendix 5, wherein the second information indicates the reliability of the second sample. (Appendix 8) The information processing apparatus according to Appendix 5, wherein the first determination unit adjusts the second sample in the joint training based on the second information received from the second determination unit. (Supplementary Note 9) The information processing apparatus according to Supplementary Note 5, wherein the second determination unit adjusts the shared class in the joint training based on the first information received from the first determination unit. (Supplementary Note 10) An information processing method for classifying a second sample having no label of a second domain different from the first domain based on a first sample having a label of the first domain, wherein the class of the second sample is a subset or the entire set of the class of the first sample. A step of determining, by a first determination unit, whether the class of the first sample is a shared class between the first domain and the second domain; A step of determining, by a second determination unit, the shared class to which the second sample belongs; A step of performing joint training on the first determination unit and the second determination unit by exchanging information. An information processing method including the above steps. (Supplementary Note 11) The information processing method according to Supplementary Note 10, wherein the first determination unit and the second determination unit are realized by a convolutional neural network model. (Supplementary Note 12) The information processing method according to Supplementary Note 11, wherein the first determination unit and the second determination unit are realized by a convolutional neural network model having the same structure. (Supplementary Note 13) The information processing method according to Supplementary Note 10, wherein the first sample and the second sample are images. (Supplementary Note 14) In the joint training, the first determination unit supplies first information regarding the shared class to the second determination unit, and the second determination unit supplies second information regarding the second sample to the first determination unit. The information processing method according to any one of Supplementary Notes 10 to 13. (Supplementary Note 15) The information processing method according to Supplementary Note 14, wherein the first information indicates the probability that the class of the first sample is the shared class. (Supplementary Note 16) The second information is the information processing method described in Supplementary Note 14 that indicates the reliability of the second sample. (Supplementary Note 17) In the collaborative training, the first determination unit adjusts the second sample based on the second information received from the second determination unit, which is the information processing method described in Supplementary Note 14. (Supplementary Note 18) In the collaborative training, the second determination unit adjusts the shared class based on the first information received from the first determination unit, which is the information processing method described in Supplementary Note 14. (Supplementary Note 19) A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a computer, it realizes the information processing method described in any one of Supplementary Notes 10 to 18.

[0092] The above describes specific embodiments of the present disclosure, but those skilled in the art can make various changes, improvements, or equivalents to the present disclosure within the gist and scope of the appended claims. These changes, improvements, or equivalents belong to the protection scope of the present disclosure.

Claims

1. An information processing apparatus for classifying a second sample having no label of a second domain different from the first domain based on a first sample having a label of the first domain, wherein a class of the second sample is a subset or the entire set of the class of the first sample. The information processing apparatus, a first determination unit that determines whether the class of the first sample is a shared class between the first domain and the second domain; a second determination unit that determines a shared class to which the second sample belongs, and includes: The information processing apparatus in which the first determination unit and the second determination unit execute joint training by exchanging information.

2. The information processing apparatus according to claim 1, wherein the first determination unit and the second determination unit are realized by a convolutional neural network model.

3. The information processing apparatus according to claim 2, wherein the first determination unit and the second determination unit are realized by a convolutional neural network model having the same structure.

4. The information processing apparatus according to claim 1, wherein the first sample and the second sample are images.

5. In the joint training, the first determination unit supplies first information regarding the shared class to the second determination unit, and the second determination unit supplies second information regarding the second sample to the first determination unit. The information processing apparatus according to any one of claims 1 to 4.

6. The information processing apparatus according to claim 5, wherein the first information indicates a probability that the class of the first sample is the shared class.

7. The information processing apparatus according to claim 5, wherein the second information indicates the reliability of the second sample.

8. The information processing apparatus according to claim 5, wherein the first determination unit adjusts the second sample in the joint training based on the second information received from the second determination unit.

9. The information processing apparatus according to claim 5, wherein the second determination unit adjusts the shared class in the joint training based on the first information received from the first determination unit.

10. An information processing method for classifying a second sample having no label of a second domain different from the first domain based on a first sample having a label of the first domain, wherein the class of the second sample is a subset or the entire set of the class of the first sample. A step in which a first determination unit determines whether the class of the first sample is a shared class between the first domain and the second domain. A step in which a second determination unit determines the shared class to which the second sample belongs. An information processing method including a step of performing collaborative training on the first determination unit and the second determination unit by exchanging information.

Citation Information

Patent Citations

  • Transfer learning in neural networks

    JP2018525734A

  • Knowledge transfer method, information processing apparatus, and storage medium

    JP2019215861A

  • Ensemble transfer learning

    US20180314975A1

  • Model generation device, model adjustment device, model generation method, model adjustment method, and recording medium

    WO2020202591A1