Sampling device and sampling method

The proposed sampling technique addresses the issue of sampling bias in data labeling for machine learning by using clustering and selective data sampling, resulting in improved accuracy and reduced sample size for image recognition models.

JP7695097B2Active Publication Date: 2025-06-18DENSO TEN LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021064088
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-05
Publication Date
2025-06-18
Estimated Expiration
2041-04-05

AI Technical Summary

Technical Problem

Existing data labeling techniques for machine learning, such as semi-supervised learning, are prone to sampling bias, which decreases the accuracy of image recognition models, and increasing the number of labeled data does not effectively suppress this bias.

Method used

A sampling technique that uses clustering to identify feature vectors, discriminates between data groups close to and far from the cluster center, and selectively samples data from these groups to minimize sampling bias while reducing the overall number of samples.

Benefits of technology

This approach effectively suppresses sampling bias and reduces the number of samples required for accurate machine learning model training, thereby improving the accuracy of image recognition tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695097000001
    Figure 0007695097000001
  • Figure 0007695097000002
    Figure 0007695097000002
  • Figure 0007695097000003
    Figure 0007695097000003
Patent Text Reader

Abstract

To provide a sampling technique capable of suppressing the occurrence of sampling bias while suppressing the number of samples.SOLUTION: The sampling device includes a clustering unit and a selection unit. The clustering unit acquires a feature vector from each of the multiple data and classifies the multiple data into multiple clusters based on the distance of the feature vector. The selection unit discriminates a first data group near the centroid of the cluster and a second data group other than the first data group concerning each of the multiple clusters and selects the data from the first data group.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for sampling data to be labeled from a large amount of data.

Background Art

[0002] For image recognition such as image classification and object detection using machine learning such as deep learning, a large number of images ranging from several thousand to over several million are required. And it is necessary to perform labeling to specify the type of object, the object detection area, etc. for the large number of images. Similarly, labeling is required for data classification and data recognition using machine learning even when the data is not an image.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In order to reduce the man-hours of this labeling work, it is conceivable to use a learning method such as semi-supervised learning that learns using both labeled data and unlabeled data. However, in this learning method, if there is a bias in the data to be labeled, sampling bias occurs. When sampling bias occurs in the labeled data, bias also occurs in the learning using the unlabeled data, and as a result, the accuracy of the learned model for image recognition decreases.

[0005] It may be considered that to suppress the occurrence of sampling bias, the number of data to be labeled may be increased. However, increasing the number of data to be labeled not only makes it difficult to reduce the man-hours of the labeling work, but also the possibility of bias in the data to be labeled still remains.

[0006] Note that the invention described in Patent Document 1 selects uncertain data as the labeling target and does not suppress the occurrence of the sampling bias described above.

[0007] In view of the above problems, an object of the present invention is to provide a sampling technique capable of suppressing the occurrence of sampling bias while suppressing the number of samples.

Means for Solving the Problems

[0008] The sampling device according to the present invention includes a clustering unit that acquires a feature vector for each of a plurality of data and classifies the plurality of data into a plurality of clusters based on the distance of the feature vectors, and for each of the plurality of clusters, a first data group that is close to the center of gravity of the cluster and a second data group that is other than the first data group are discriminated, and a selection unit that selects the data from among the first data group. (First configuration).

[0009] In the sampling device having the first configuration, the selection unit may be configured to select the data closest to the center of gravity of the cluster (second configuration).

[0010] In the sampling device having the first or second configuration, the selection unit divides the second data group into a third data group that is far from the center of gravity of the cluster and a fourth data group that is other than the third data group, and may be configured to select the data from among the third data group (third configuration).

[0011] In the sampling device having the third configuration, the selection unit preferentially selects the data that is farther from the center of gravity of the cluster from among the third data group to (Fourth configuration).

[0012] In the sampling device having any one of the first to fourth configurations described above, the selection unit may be configured (a fifth configuration) to increase the number of data selected from each of the plurality of clusters as the variance of the feature vectors within the cluster is larger.

[0013] In the sampling device having any one of the first to fourth configurations described above, the selection unit may be configured (a sixth configuration) to increase the number of data selected from each of the plurality of clusters as the number of data within the cluster is larger.

[0014] The sampling method according to the present invention includes a clustering step of obtaining a feature vector for each of a plurality of data and classifying the plurality of data into a plurality of clusters based on the distance of the feature vectors, and a selection step of discriminating, for each of the plurality of clusters, a first data group close to the center of gravity of the cluster and a second data group other than the first data group, and selecting the data from among the first data group (a seventh configuration).

Advantages of the Invention

[0015] According to the present invention, it is possible to suppress the occurrence of sampling bias while suppressing the number of samples.

Brief Description of the Drawings

[0016]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiments for Carrying Out the Invention

[0017] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings.

[0018] <1. Configuration of Information Processing Apparatus> FIG. 1 is a diagram showing a schematic configuration example of an information processing apparatus according to an embodiment. The information processing apparatus 1 is an example of a sampling apparatus. The information processing apparatus 1 may be an information processing apparatus installed in a single location, or may be a distributed information processing apparatus in which components are installed in a plurality of locations in a distributed manner.

[0019] The information processing apparatus 1 includes a control unit 11 and a storage unit 12.

[0020] The control unit 11 is a computer including at least one processor. Specifically, the control unit 11 is a computer including a CPU (Central Processing Unit), a RAM (Random Access Memory), and a ROM (Read Only Memory) (not shown). The control unit 11 processes and transmits / receives information based on a program stored in the storage unit 12, and controls the entire information processing apparatus 1.

[0021] The control unit 11 includes a clustering unit 11a and a selection unit 11b. Various functions of the control unit 11 such as the clustering unit 11a are realized by the CPU executing arithmetic processing according to a program stored in the storage unit 12.

[0022] The clustering unit 11a acquires a feature vector for each of a plurality of data, and classifies the plurality of data into a plurality of clusters based on the distance of the feature vectors.

[0023] The selection unit 11b discriminates, for each of the plurality of clusters, a first data group close to the centroid of the cluster and a second data group other than the first data group, and selects data from among the first data group.

[0024] <2. Operation of the information processing apparatus> FIG. 2 is a flowchart showing an operation example of the information processing apparatus 1. When the information processing apparatus 1 is powered on and a plurality of data are input, the information processing apparatus 1 starts the operation of the flowchart shown in FIG. 2. The input form of the plurality of data is not particularly limited. The information processing apparatus 1 may receive the plurality of data by wireless communication, may receive the plurality of data by wired communication, or may input the plurality of data via a storage medium detachable from the information processing apparatus 1. The plurality of data input to the information processing apparatus 1 is composed of, for example, thousands to over millions of images. Note that the plurality of data input to the information processing apparatus 1 may be data other than images.

[0025] First, the clustering unit 11a acquires a feature vector for each of the plurality of data (step S10). The acquisition of the feature vector is performed by, for example, machine learning. This machine learning may be supervised machine learning or unsupervised machine learning. Also, the number of dimensions of the feature vector is not particularly limited. In the following description, for the sake of simplicity of explanation and figures, etc., the case where the feature vector is a two-dimensional vector will be taken as an example. As unsupervised machine learning, for example, DNN (Deep Neural Network) or the like can be used.

[0026] Next, the clustering unit 11a classifies the plurality of data into a plurality of clusters based on the distance of the feature vectors (step S20). The clustering unit 11a performs, for example, hierarchical clustering and classifies the plurality of data into a plurality of clusters by fixing the number of finally classified clusters. As the distance of the feature vectors, for example, Euclidean distance, cosine similarity, etc. can be used.

[0027] FIG. 3 is a diagram schematically showing an example of the result of classifying a plurality of data into first to fourth clusters CL1 to CL4 by the clustering unit 11a. Each black circle in FIG. 3 represents individual data, but only a smaller number than the actual number of data is shown for simplicity of explanation. Hereinafter, the description will continue using the example shown in FIG. 3.

[0028] Next, for each of the plurality of clusters, the selection unit 11b discriminates between a first data group close to the centroid of the cluster and a second data group other than the first data group, and selects the data from among the first data group (step S30). As a result, typical data of each cluster is selected, so that it is possible to suppress the occurrence of sampling bias while suppressing the number of samples. Note that the centroid of a cluster means the centroid of the feature vector of the data belonging to the cluster.

[0029] FIG. 4 is a diagram schematically showing the result of discriminating between a first data group G1 close to the centroid C1 of the first cluster CL1 and a second data group G2 other than the first data group G1 for the first cluster CL1. In FIG. 4, the first data group G1 includes three pieces of data, but the method of determining the boundary between the first data group G1 and the second data group G2 is not particularly limited. The number of data belonging to the first data group G1 may be determined in consideration of, for example, the number of clusters, the number of data belonging to the cluster, and the number of data selected by the selection unit 11b. When determining the number of data belonging to the first data group G1 based on the number of data belonging to the cluster, for example, the number of data in the first data group G1 is 20% of the number of data belonging to the cluster, and the number of data in the second data group G2 is 80% of the number of data belonging to the cluster.

[0030] Further, the selection unit 11b may use the data belonging to the first data group G1 as data within a first predetermined distance from the centroid C1 of the cluster. Also, the value of the first predetermined distance may be determined in consideration of the number of clusters and the number of data selected by the selection unit 11b. By determining in this way, it is possible to realize the selection of typical data of the cluster and the suppression of sampling bias by selecting a plurality of data.

[0031] In this embodiment, the selection unit 11b selects the data closest to the centroid of the cluster. That is, since the most typical data of each cluster is selected, it is possible to further suppress the occurrence of sampling bias while suppressing the number of samples.

[0032] Specifically, using the example shown in FIG. 3, in this embodiment, the selection unit 11b preferentially selects the data D1 (see FIG. 4) closest to the centroid C1 of the first cluster CL1 from among the first cluster CL1, and preferentially selects the data closest to the centroid C2 of the second cluster CL2 from among the second cluster CL2, preferentially selects the data closest to the centroid C3 of the third cluster CL3 from among the third cluster CL3, and preferentially selects the data closest to the centroid C4 of the fourth cluster CL4 from among the fourth cluster CL4.

[0033] Referring to FIGS. 2 and 5, further, the selection unit 11b divides the second data group G2 into a third data group G3 far from the centroid of the cluster and a fourth data group G4 close to the centroid of the cluster. Also, the selection unit 11b selects data from among the third data group (step S30). The division between the third data group G3 and the fourth data group G4 may be performed based on the distance from the centroid of the cluster or the number of data. Thereby, since the data near the boundary of each cluster (the portion farthest from the centroid of each cluster) is selected, it is possible to suppress the occurrence of sampling bias while efficiently suppressing the number of samples.

[0034] Using FIG. 5, the division between the third data group G3 and the fourth data group G4 will be described. FIG. 5 is a diagram schematically showing the result of dividing the second data group G2 into a third data group G3 far from the centroid C1 of the first cluster CL1 and a fourth data group G4 other than the third data group G3 with respect to the first cluster CL1. The method of determining the boundary between the third data group G3 and the fourth data group G4 is not particularly limited. For example, considering the number of clusters, the number of data belonging to the second data group G2, and the number of data selected by the selection unit 11b, the number of data belonging to the third data group G3 may be determined. When determining the number of data belonging to the third data group G3 based on the number of data belonging to the second data group G2, for example, the number of data in the third data group G3 is 50% of the number of data belonging to the second data group G2. Also, for example, the data belonging to the third data group G3 may be data at a second predetermined distance or more from the centroid of the cluster, and the value of the second predetermined distance may be determined in consideration of the number of clusters and the number of data selected by the selection unit 11b.

[0035] In the present embodiment, the selection unit 11b preferentially selects data farther from the centroid of the cluster from among the third data group. That is, among the special (atypical) data of each cluster, data with a higher degree of specialness is preferentially selected, so that the occurrence of sampling bias can be further suppressed while suppressing the number of samples. Note that, unlike the present embodiment, for example, the selection unit 11b may preferentially select data closer to the centroid of the cluster from among the third data group, and also, for example, the selection unit 11b may randomly select a predetermined number of data from among the third data group.

[0036] The data selection method will be specifically described using the examples shown in FIGS. 3 and 5. In the present embodiment, the selection unit 11b selects at least one data from among the first data group G1 (the data group in the region closest to the center of gravity C1 of the first cluster CL1) of the first cluster CL1. Examples of the data selection method from among the first data group G1 include: (1) a method of preferentially selecting data with a shorter distance from the center of gravity C1 of the first cluster CL1; (2) a method of preferentially selecting data with a longer distance from the center of gravity C1 of the first cluster CL1; and (3) a method of randomly selecting from among the first data group G1. Similarly, for the second to fourth clusters CL2 to CL4, data is selected from the data group in the region closest to the center of gravity of the cluster.

[0037] In the present embodiment, when the data selection from the data group in the region closest to the center of gravity of the cluster is completed, the selection unit 11b selects at least one data from among the third data group G3 (the data group in the region farthest from the center of gravity C1 of the first cluster CL1) of the first cluster CL1. Examples of the data selection method from among the third data group G3 include: (1) a method of preferentially selecting data with a longer distance from the center of gravity C1 of the first cluster CL1; (2) a method of preferentially selecting data with a shorter distance from the center of gravity C1 of the first cluster CL1; and (3) a method of randomly selecting from among the third data group G3. Similarly, for the second to fourth clusters CL2 to CL4, data is selected from the data group in the region farthest from the center of gravity of the cluster.

[0038] In this embodiment, even if the selection of data from the data group in the region closest to the centroid of the cluster and the data group 3 in the region farthest from the centroid of the cluster is completed and the number of data selected by the selection unit 11b does not reach a preset value, the selection unit 11b selects at least one data from the fourth data group G4 (the data group in the region second farthest from the centroid C1 of the first cluster CL1) of the first cluster CL1. Examples of the method for selecting data from the fourth data group G4 include: (1) a method of preferentially selecting data with a greater distance from the centroid C1 of the first cluster CL1; (2) a method of preferentially selecting data with a shorter distance from the centroid C1 of the first cluster CL1; and (3) a method of randomly selecting from the fourth data group G4. Similarly, for the second to fourth clusters CL2 to CL4, data is selected from the data group in the region second farthest from the centroid of the cluster.

[0039] For example, the selection unit 11b first selects the data in the region closest to the centroid of the cluster, second selects the data in the region farthest from the centroid of the cluster, and third selects the data in the region second farthest from the centroid of the cluster. By selecting in this way, it is possible to achieve both suppression of the number of samples and suppression of sampling bias.

[0040] When the number of data selected by the selection unit 11b reaches a preset value, the selection unit 11b ends the selection process. For example, if the number of data selected by the selection unit 11b is set to 100 and a plurality of data are classified into the first to fourth clusters CL1 to CL4 as in the example shown in FIG. 3, the selection unit 11b selects 25 data from each of the first to fourth clusters CL1 to CL4.

[0041] When the process of step S30 ends and the data selected by the selection unit 11b is output, the information processing apparatus 1 ends the operation of the flowchart shown in FIG. 2. The output form of the data selected by the selection unit 11b is not particularly limited. The information processing apparatus 1 may transmit the data selected by the selection unit 11b by wireless communication, may transmit the data selected by the selection unit 11b by wired communication, or may output the data selected by the selection unit 11b via a storage medium detachable from the information processing apparatus 1.

[0042] <3. Modification Example> The above-described embodiments should be considered to be illustrative in all respects and not restrictive. The technical scope of the present invention is shown not by the description of the above embodiments but by the scope of the claims, and it should be understood that all modifications belonging to the meaning and scope equivalent to the scope of the claims are included.

[0043] For example, the selection unit 11b may increase the number of pieces of the data selected from each of the plurality of clusters as the variance of the feature amount vectors within the cluster is larger. Thereby, it can be expected to efficiently suppress the occurrence of sampling bias in each cluster. For example, according to the variances V1 to V4 of the feature amount vectors in the first to fourth clusters CL1 to CL4, the selection unit 11b may select data such that the number of data selected from the first cluster CL1: the number of data selected from the second cluster CL2: the number of data selected from the third cluster CL3: the number of data selected from the fourth cluster CL4 = V1: V2: V3: V4.

[0044] For example, the selection unit 11b may select a larger number of data from each of the plurality of clusters as the number of data within the cluster is larger. Thereby, it can be expected to efficiently suppress the occurrence of sampling bias in each individual cluster. For example, according to the number of data N1 to N4 in the first to fourth clusters CL1 to CL4, the selection unit 11b may select data such that the number of data selected from the first cluster CL1: the number of data selected from the second cluster CL2: the number of data selected from the third cluster CL3: the number of data selected from the fourth cluster CL4 = N1: N2: N3: N4.

[0045] In the above-described embodiment, an example in which the control unit 11 and the storage unit 12 have separate configurations has been described, but the present invention is not limited thereto. The storage unit 12 may be included in the control unit 11. Further, the storage unit 12 may be a storage medium that can be attached to and detached from the control unit 11. Here, the recording medium is, for example, a flexible disk, a hard disk, a CD-ROM, an MO, a DVD, a DVD-ROM, a DVD-RAM, a large-capacity DVD, a next-generation DVD, or a semiconductor memory.

Explanation of Reference Numerals

[0046] 1 Information processing apparatus 11 Control unit 11a Clustering unit 11b Selection unit 12 Storage unit

Claims

1. For each of a plurality of data, a feature vector is obtained, and a clustering unit that classifies the plurality of data into a plurality of clusters based on the distance of the feature vectors, For each of the plurality of clusters, a first data group close to the centroid of the cluster and a second data group other than the first data group are discriminated, the data is selected from among the first data group, and for each of the plurality of clusters, the greater the variance of the feature vectors within the cluster, the greater the number of the data selected from the cluster. A selection unit; A sampling device comprising:

2. The sampling device according to claim 1, wherein the selection unit selects the data closest to the centroid of the cluster.

3. The sampling device according to claim 1 or claim 2, wherein the selection unit divides the second data group into a third data group far from the centroid of the cluster and a fourth data group close to the centroid of the cluster, and selects the data from among the third data group.

4. The sampling device according to claim 3, wherein the selection unit preferentially selects the data farther from the centroid of the cluster from among the third data group.

5. The sampling device according to any one of claims 1 to 4, wherein the selection unit increases the number of the data selected from the cluster as the number of the data within the cluster increases for each of the plurality of clusters.

6. A clustering step of obtaining a feature vector for each of a plurality of data and classifying the plurality of data into a plurality of clusters based on the distance of the feature vectors, For each of the plurality of clusters, discriminate a first data group close to the center of gravity of the cluster and a second data group other than the first data group, select the data from among the first data group, and for each of the plurality of clusters, a selection step of increasing the number of the data selected from the cluster as the variance of the feature vectors in the cluster is larger; A sampling method comprising the above.

Citation Information

Patent Citations

  • Method and device for automatically acquiring multi-level classification training data of enterprise

    CN112287075A

  • Pattern learning method and device

    JP1997034862A

  • Characteristic estimation device

    JP2012208710A

  • Sensor node, server apparatus, and identification system, method and program

    JP2020119238A

  • Physical property prediction device and physical property prediction method

    JP2020187417A