Expression learning device, method, and program
The expression learning device enhances clustering accuracy by distinguishing between target and non-target features through a multi-unit approach, addressing the challenge of accurately representing complex data in machine learning.
Patent Information
- Application Number
- JP2022147323
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-09-15
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2042-09-15
AI Technical Summary
Existing representation learning methods struggle to accurately distinguish between features of interest and non-interest in complex data, leading to suboptimal clustering and classification performance.
An expression learning device and method that utilizes a first and second acquisition unit, vector calculation units, a similarity calculation unit, and an update unit to calculate and update model parameters based on similarities and loss functions, focusing on both target and non-target features to enhance clustering accuracy.
The solution enables highly accurate clustering and classification by emphasizing features of interest and suppressing non-interest features, improving clustering accuracy and separation in latent spaces.
Smart Images

Figure 0007815072000008 
Figure 0007815072000009 
Figure 0007815072000010
Abstract
Description
[Technical Field]
[0001] FIELD Embodiments of the present invention relate to an expression learning device, method, and program. [Background technology]
[0002] In recent machine learning, a representation learning method has been proposed that represents complex data such as images, audio, and time series data as low-dimensional feature vectors. As an example, a representation learning method suitable for clustering has been proposed. This method learns features that cluster complex and abstract information that can be grouped, so it learns features that cluster both features that you want to pay attention to and features that you don't want to pay attention to. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] “Clustering Friendly representation learning via instance discrimination and feature decorrelation”, Yaling Tao, Kentaro Takagi, Kouta Nakata. arXiv:2106.00131 (ICLR2021) Summary of the Invention [Problem to be solved by the invention]
[0004] The problem to be solved by the present invention is to provide an expression learning device, method, and program that realizes highly accurate expression learning. [Means for solving the problem]
[0005] An expression learning device according to an embodiment includes a first acquisition unit, a second acquisition unit, a first vector calculation unit, a second vector calculation unit, a similarity calculation unit, a loss function calculation unit, and an update unit. The first acquisition unit acquires target data. The second acquisition unit acquires non-target data similar to a non-target feature included in the target data. The first vector calculation unit calculates a latent vector in the latent space of the target data using first model parameters related to a first machine learning model to be trained. The second vector calculation unit calculates a first non-target latent vector in the latent space of a non-target feature in the target data and a second non-target latent vector in the latent space of the non-target data using second model parameters related to a second machine learning model to be trained. The similarity calculation unit calculates a first similarity by correcting the similarity between the latent vector and a first representative value of the latent vector by the similarity between the first non-interest latent vector and a second representative value of the first non-interest latent vector, and calculates a second similarity between the second non-interest latent vector and a third representative value of the second non-interest latent vector. The loss function calculation unit calculates a loss function including the first similarity and the second similarity. The update unit updates the first model parameters and / or the second model parameters based on the loss function. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of an expression learning device according to an embodiment of the present invention; [Figure 2] A diagram showing an example of target data [Figure 3] Figure showing other examples of target data [Figure 4] FIG. 10 is a diagram showing an example of a processing procedure for expression learning processing according to the present embodiment; [Figure 5] FIG. 5 shows various components and data flows involved in the expression learning process shown in FIG. [Figure 6] FIG. 6 is a diagram showing an example of target data used in the expression learning process illustrated in FIGS. 4 and 5. [Figure 7] FIG. 7 is a diagram showing an example of non-interest data corresponding to FIG. 6; [Figure 8]FIG. 10 is a diagram showing the clustering processing procedure according to the first example of use. [Figure 9] FIG. 10 is a diagram showing the clustering accuracy of this embodiment and a comparative example. [Figure 10] FIG. 10 is a diagram showing the results of visualizing the data distribution in the latent space by compressing it into two dimensions according to the embodiment and the comparative example. [Figure 11] FIG. 10 is a diagram showing the clustering processing procedure according to the second example of use. [Figure 12] FIG. 10 is a diagram showing the clustering processing procedure according to the third example of use. [Figure 13] FIG. 10 is a diagram showing the clustering processing procedure according to the fourth example of use. [Figure 14] FIG. 10 is a diagram showing the procedure of search processing according to Usage Example 5. [Figure 15] FIG. 10 is a diagram showing the procedure of search processing according to Usage Example 6. [Figure 16] FIG. 10 is a diagram showing the procedure of search processing according to Usage Example 7. [Figure 17] FIG. 10 is a diagram showing the processing procedure of search processing according to Usage Example 8. DETAILED DESCRIPTION OF THE INVENTION
[0007] Hereinafter, an expression learning device, method, and program according to this embodiment will be described with reference to the drawings.
[0008] Fig. 1 is a diagram showing an example of the configuration of an expression learning device 100 according to this embodiment. As shown in Fig. 1, the expression learning device 100 is a computer having a processing circuit 1, a storage device 2, an input device 3, a communication device 4, and a display device 5. Data communication between the processing circuit 1, the storage device 2, the input device 3, the communication device 4, and the display device 5 is performed via a bus.
[0009] The processing circuit 1 includes a processor such as a CPU (Central Processing Unit) and a memory such as a RAM (Random Access Memory). The processing circuit 1 includes a first acquisition unit 11, a second acquisition unit 12, a first vector calculation unit 13, a second vector calculation unit 14, a similarity calculation unit 15, a loss function calculation unit 16, an update unit 17, a learning control unit 18, a post-processing unit 19, and a display control unit 20. The processing circuit 1 executes an expression learning program to realize the functions of the above-mentioned units 11 to 20. The expression learning program is stored in a non-transitory computer-readable recording medium such as a storage device 2. The expression learning program may be implemented as a single program that describes all the functions of the above-mentioned units 11 to 20, or may be implemented as multiple modules divided into several functional units. Furthermore, the above-mentioned units 11 to 20 may be implemented using integrated circuits such as application-specific integrated circuits (ASICs). In this case, the units may be implemented on a single integrated circuit or individually on multiple integrated circuits.
[0010] The first acquisition unit 11 acquires learning data to be processed (hereinafter, target data). The target data refers to data to be classified by a machine learning model. The target data has features to be focused on (hereinafter, target features) and features not to be focused on (hereinafter, non-target features). The target data is not particularly limited as long as it can be classified, and for example, image data, audio data, character data, waveform data, etc. may be used.
[0011] Here, a specific example of the target data will be described with reference to FIGS.
[0012] Fig. 2 is a diagram showing an example of target data. The target data shown in Fig. 2 are optical images 201-212 depicting birds (hereinafter referred to as bird images). The bird images 201-212 are part of a data set called Birds400. An example of a feature of interest in a bird image is the bird part depicted in each of the images 201-212. An example of a non-interest feature is the background, such as the sky, trees, or ground depicted in each of the images 201-212.
[0013] Fig. 3 is a diagram showing another example of target data. The target data shown in Fig. 3 are optical images 301 to 312 depicting people's faces (hereinafter referred to as "person images"). Examples of features of interest in the person images are the hair color, beard, and mouth depicted in each of the images 301 to 312. Examples of features of non-interest are glasses, hats, etc. depicted in each of the images 301 to 312.
[0014] The second acquisition unit 12 acquires non-interest data that is data similar to the non-interest features included in the target data. The non-interest data b is not particularly limited as long as it can be classified, similar to the target data, and for example, image data, audio data, character data, waveform data, etc. may be used.
[0015] The first vector calculation unit 13 calculates a latent vector in the latent space of the target data using the first model parameters of the first machine learning model of the training target. The latent vector is a vector representing data in which the dimensions of the target data are compressed. The latent space refers to a space spanned by the latent vectors. The first machine learning model is an encoder network that converts the target data into a latent vector. The first model parameters are parameters of the training target that are assigned to the first machine learning model. Typically, the first model parameters are weights and biases. The first model parameters are stored in the storage device 2.
[0016] The second vector calculation unit 14 uses second model parameters of the second machine learning model to be trained to calculate a latent vector in the latent space of a feature of interest in the target data (hereinafter referred to as a first latent vector of interest). The second vector calculation unit 14 also uses the second model parameters to calculate a second latent vector of interest in the latent space of the data of interest. Hereinafter, when there is no need to distinguish between the first latent vector of interest and the second latent vector of interest, they will simply be referred to as a latent vector of interest. The latent space related to the second machine learning model refers to a space spanned by latent vectors of interest. The second machine learning model is an encoder network that converts features of interest or data of interest in the target data into a first latent vector of interest and a second latent vector of interest. The second model parameters are parameters of the training target assigned to the second machine learning model. Typically, the second model parameters are weights and biases. The second model parameters are stored in the storage device 2. The first machine learning model and the second machine learning model may be of the same type or may be of different types.
[0017] The similarity calculation unit 15 calculates a first similarity by correcting the similarity between the latent vector and the first representative value of the latent vectors by the similarity between the first non-interest latent vector and the second representative value of the first non-interest latent vector. The first representative value is a value representative of the multiple latent vectors obtained up to the previous iteration of the expression learning process. Similarly, the second representative value is a value representative of the multiple first non-interest latent vectors obtained up to the previous iteration of the expression learning process. Furthermore, the similarity calculation unit 15 calculates a second similarity between the second non-interest latent vector and a third representative value of the second non-interest latent vector. The third representative value is a value representative of the multiple second non-interest latent vectors obtained up to the previous iteration of the expression learning process. Furthermore, the similarity calculation unit 15 may calculate a third similarity between the first non-interest latent vector and the third representative value.
[0018] The loss function calculation unit 16 calculates a loss function including at least the first similarity and the second similarity. When the third similarity is calculated, the loss function calculation unit 16 may calculate a loss function including the first similarity, the second similarity, and the third similarity.
[0019] The update unit 17 updates the first model parameters and / or the second model parameters based on the loss function. More specifically, the update unit 17 updates the first model parameters and / or the second model parameters according to the gradient of the loss function.
[0020] The learning control unit 18 controls the expression learning process. Specifically, the learning control unit 18 determines whether a stop condition for the expression learning process is satisfied, and repeats the expression learning process until it is determined that the stop condition is satisfied. If it is determined that the stop condition for the expression learning process is satisfied, the learning control unit 18 outputs the first model parameters and / or the second model parameters for the current iteration count as learned model parameters.
[0021] The post-processing unit 19 performs post-processing utilizing the information resources obtained by the representation learning process. Specifically, the post-processing includes clustering and search. The clustering process clusters the target data or new data using the latent vectors and / or non-focus latent vectors obtained by the representation learning process. The search process searches for other target data or new data similar to the reference target data or new data using the latent vectors and / or non-focus latent vectors obtained by the representation learning process.
[0022] The display control unit 20 displays various data on the display device 5. For example, the display control unit 20 displays the results of clustering using a machine learning model.
[0023] The storage device 2 is configured with a ROM (Read Only Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), an integrated circuit storage device, etc. The storage device 2 stores an expression learning program, etc. The storage device 2 also stores the latent vector of the target data, its first representative value, the non-interest latent vector of the target data, its second representative value, the non-interest latent vector of the non-interest data, and its third representative value.
[0024] The input device 3 inputs various commands from a user. Examples of the input device 3 that can be used include a keyboard, a mouse, various switches, a touchpad, and a touch panel display. An output signal from the input device 3 is supplied to the processing circuit 1. Note that the input device 3 may also be an input device of a computer connected to the processing circuit 1 via a wired or wireless connection.
[0025] The communication device 4 is an interface for performing data communication between the expression learning device 100 and an external device connected to the expression learning device 100 via a network.
[0026] The display device 5 displays various information. For example, the display device 5 displays various data under the control of the output control unit 16. As the display device 5, a CRT (Cathode-Ray Tube) display, a liquid crystal display, an organic EL (Electro Luminescence) display, an LED (Light-Emitting Diode) display, a plasma display, or any other display known in the art can be used as appropriate. The display device 5 may also be a projector.
[0027] The expression learning process according to this embodiment will be described below.
[0028] FIG. 4 is a diagram showing an example of the processing procedure of the expression learning process according to this embodiment. FIG. 5 is a diagram showing various components and data flows related to the expression learning process shown in FIG. 4. The Sx memory 21 shown in FIG. 5 is part of the storage device 2 and is a memory that stores latent vectors and first representative values. The Zx memory 22 is part of the storage device 2 and is a memory that stores first non-interest latent vectors and second representative values. The Zb memory 23 is part of the storage device 2 and is a memory that stores second non-interest latent vectors and third representative values. The similarity calculation module 151 is part of the similarity calculation unit 15 and calculates a first similarity. The similarity calculation module 152 is part of the similarity calculation unit 15 and calculates a second similarity. The similarity calculation module 153 is part of the similarity calculation unit 15 and calculates a third similarity.
[0029] 4 and 5, the first acquisition unit 11 acquires target data x (step S401). In step S401, the first acquisition unit 11 acquires N pieces of target data x i N is a natural number equal to or greater than 2, and the specific value can be set arbitrarily. N pieces of target data x i and M pieces of non-interest data b i constitutes one mini-batch. The subscript i represents the i-th target data x and non-interest data b. Note that N and M may be the same number or different numbers.
[0030] FIG. 6 is a diagram showing an example of target data used in the expression learning process exemplified in FIGS. 4 and 5. As shown in FIG. 6, each target data is an image in which numbers from 0 to 9 are drawn against a background of stripes at various angles. In FIG. 6, l represents the value of the number drawn in the image, i.e., the label, and bg represents the angle of the stripes. For example, the image in the upper left has l=5 and bg=90, meaning that the image is an image in which the number 5 is drawn with 90-degree stripes. Here, the number is a feature of interest, and the stripes are a non-interest feature. The target data corresponds, for example, to a defect inspection image in which a defective product is drawn.
[0031] In step S401, the first acquisition unit 11 acquires the target data x i Data extension may be performed on the image data by, for example, randomly cropping the image or randomly changing the brightness, luminance, saturation, etc.
[0032] When step S401 is performed, the first vector calculation unit 13 calculates the target data x using the first model parameters. i Latent vector Sx i (Step S402). The first vector calculation unit 13 reads out the first model parameters of the training target from the storage device 2, sets the read out first model parameters to the first machine learning model, and calculates the target data x i By forward propagating, the latent vector Sx i The first machine learning model calculates the d-dimensional target data x i Enter the d'-dimensional latent vector Sx i The encoder network outputs d', where d' is smaller than d. The architecture of the encoder network is not particularly limited, but a deep neural network such as ResNet (Deep Residual Learning) may be used. The model parameters in the first update of the representation learning process may be set to any value. The latent vector Sx i is stored in the Sx memory 21.
[0033] When step S402 is performed, the second vector calculation unit 14 calculates the target data x using the second model parameters. i The non-focused latent vector Zx i In step S403, the second vector calculation unit 14 reads out the second model parameters of the training target from the storage device 2, sets the read out second model parameters in the second machine learning model, and assigns the target data x i By forward propagating, the non-focused latent vector Zx i The second machine learning model calculates the d-dimensional target data xi Enter the d´´-dimensional non-focus latent vector Zx i The encoder network outputs the non-focus latent vector Zx. d' is smaller than d. The model parameters in the first update of the representation learning process can be set to any value. i is stored in the Zx memory 22.
[0034] When step S403 is performed, the similarity calculation unit 15 calculates the latent vector Sx i and its representative value S´x j and the non-focus latent vector Zx i and its representative value Z´x j Based on this, the similarity S1 ij (Step S404). The subscript "j" represents the j-th representative value S'x or Z'x. In step S404, the similarity calculation unit 15 calculates the representative value S'x from the Sx memory 21. j Obtain the representative value S'x j is a vector that represents the latent vector Sx calculated up to the previous iteration. j is, for example, the moving average value of the multiple latent vectors Sx calculated up to the previous iteration. Similarly, the similarity calculation unit 15 calculates the representative value Z'x from the Zx memory 22. j Obtain the representative value Z'x j is the number of non-focused latent vectors Zx calculated up to the previous iteration i is a vector that represents the representative value Z'x j For example, the multiple non-interest latent vectors Zx i is the moving average value of
[0035] In step S404, the similarity calculation unit 15 calculates the latent vector Sx i and the latent vector Sx i Representative value S'x j The similarity between the non-focus latent vector Zx and the non-focus latent vector Zx i Representative value Z'x j Similarity S1 corrected by the similarity with ij As an example, the similarity S1 ijis calculated according to the following formula (1): ij The numerator of is the potential vector Sx i and the representative value S'x j The denominator represents the similarity between the non-target latent vector Zx and the representative value Z´x. j τ is a parameter that controls the degree of emphasis on the similarity.
[0036]
number
[0037] Similarity S1 ij is the latent vector Sx i and its representative value S´x j The similarity between the non-focused latent vector Zx i and its representative value Z´x j The calculation method is not limited to formula (1), as long as it can be corrected by the similarity with the following formula (2). ij τ′ is a parameter that controls the degree of emphasis of the non-interest latent vector similarity.
[0038]
number
[0039] When step S404 is performed, the second acquisition unit 12 acquires the non-interest data b i In step S405, the second acquisition unit 12 acquires M pieces of non-interest data b i In step S405, the second acquisition unit 12 acquires the non-interest data b i The data extension may be the same process as in step S401 or a different process.
[0040] FIG. 7 is a diagram showing an example of non-interest data corresponding to FIG. 6. As shown in FIG. 7, each non-interest data is an image in which striped patterns of various angles are drawn. The striped patterns are non-interest features. The non-interest data can be said to be data similar to the non-interest features of the target data. Here, the non-interest data does not need to be paired with the target data. The non-interest data corresponds, for example, to a defect inspection image of a non-defective product in which no defective product is drawn.
[0041] After step S405 is performed, the second vector calculation unit 14 uses the second model parameters to calculate a non-interest latent vector Zb of the non-interest data b (step S406). The second vector calculation unit 14 reads the second model parameters of the training target from the storage device 2, sets the read second model parameters in the second machine learning model, and assigns the non-interest data b i By forward propagating the non-focused latent vector Zb i Calculate the non-focus latent vector Zb i is stored in the Zb memory 23. In step S406, the second vector calculation unit 14 uses the same model parameters as the second model parameters used in step S403.
[0042] When step S406 is performed, the similarity calculation unit 15 calculates the non-interest latent vector Zb i and its representative value Z´b j Based on this, similarity S2 ij (step S407). In step S407, the similarity calculation unit 15 acquires the representative value Z'b from the Zb memory 23. i is the non-focused latent vector Zb calculated up to the previous iteration i is a vector that represents the representative value Z´b i For example, the multiple non-focus latent vectors Zb i The moving average value of the similarity S2 ij is calculated according to the following formula (3): Similarity S2 ij The numerator of is the non-interest latent vector Zb i and the representative value Z´bj The denominator τ is a parameter that controls the degree of emphasis on the similarity.
[0043]
number
[0044] When step S407 is performed, the similarity calculation unit 15 calculates the non-interest latent vector Zx i and the non-focus latent vector Zb j Representative value Z´b j Similarity with S3 ij In step S408, the similarity calculation unit 15 calculates the representative value Z′b j Obtain the similarity S3 ij is calculated according to the following formula (4): i Z´b j is the non-focus latent vector Zx i and the representative value Z´b of the non-focused latent vector Zb j The denominator τ is a parameter that controls the degree of emphasis on the similarity.
[0045]
number
[0046] After step S408 is performed, the loss function calculation unit 16 calculates a loss function (loss) including similarities S1, S2, and S3 (step S409). The loss function (loss) is calculated, for example, according to the following formula (5). As shown in formula (5), the loss function (loss) is defined as the sum of similarities S1, S2, and S3. The first term in formula (5) corresponds to the similarity S1. The first term is the latent vector Sx of the target data x. i But, myself S´x i The similarity to S'x is high. j The lower the similarity with the target data x, the smaller the value becomes. i and Z'x jThe degree of similarity is corrected by the similarity of the non-interest data b. The second term in equation (5) corresponds to the similarity S2. The first term corrects the loss function loss so that the non-interest latent vector Zx is not included in the latent vector Sx. The second term corrects the non-interest latent vector Zb i is Z´b i The similarity to Z´b is high. j The lower the similarity between the non-focused latent vectors Zb and the target data x, the smaller the value becomes. The second term brings similar non-focused latent vectors Zb closer to each other and separates dissimilar non-focused latent vectors Zb. The third term in equation 5) corresponds to the similarity S3. The third term is the non-focused latent vector Zx of the target data x. i is the most similar non-focused latent vector Zb of non-focused data b k The value becomes smaller as the similarity to Zx increases and the similarity to other vectors decreases. The third term plays a role in moving the non-focused latent vector Zx closer to the similar non-focused latent vector Zb and away from the dissimilar non-focused latent vector Zb.
[0047]
number
[0048] The loss function loss is not limited to equation (5), and other terms may be added as shown in equation (6). The fourth, fifth, and sixth terms in equation (6) represent the feature decorrelation terms in Non-Patent Document 1. The feature decorrelation term L fd Sx represents the degree to which the latent vectors Sx are orthogonal to each other. The feature decorrelation term L fd Zx represents the degree to which non-interest latent vectors Zx are orthogonal to each other. The feature decorrelation term L fd Zb represents the degree to which non-interest latent vectors Zb are orthogonal to each other.
[0049]
number
[0050] After step S409 is performed, the update unit 17 updates the first model parameters and / or the second model parameters according to the gradient of the loss function (step S410). In step S410, the update unit 17 can update the model parameters by using any optimization method such as stochastic gradient descent or ADAM.
[0051] When step S410 is performed, the first vector calculation unit 13 updates the representative value S'x stored in the Sx memory 21 (step S411). In step S411, the first vector calculation unit 13 updates the representative value S'x based on the latent vector Sx calculated in step S402 for the current update count. Typically, the representative value S'x is updated using a technique such as a moving average. As an example, when updating using the exponential moving average method, the latent vector Sx and the representative value S'x for the current update count are calculated according to the exponential moving average method shown in the following equation (7): old The updated representative value S'x based on new is calculated. Representative value S'x new is stored in the Sx memory 21.
[0052]
number
[0053] The update method is not limited to the above method. For example, the statistical value of the average value of the latent vector Sx calculated in step S402 for the current update count may be used as the representative value S'x for the current update count. old By replacing it with, the updated representative value S'x new may be obtained.
[0054] When step S411 is performed, the second vector calculation unit 14 updates the representative value Z'x stored in the Zx memory 22 (step S412). The updated representative value Z'x is stored in the Zx memory 22. In step S412, the second vector calculation unit 14 updates the representative value Z'x based on the non-interest latent vector Zx calculated in step S403 for the current update count. The representative value Z'x may be updated by the moving average method or the like similar to step S411.
[0055] When step S412 is performed, the second vector calculation unit 14 updates the representative value Z'b stored in the Zb memory 23 (step S413). The updated representative value Z'b is stored in the Zb memory 23. In step S413, the second vector calculation unit 14 updates the representative value Z'b based on the non-interest latent vector Zb calculated in step S406 for the current update count. The representative value Z'b may be updated by the moving average method or the like similar to step S411.
[0056] After step S413, the learning control unit 18 determines whether a stopping condition is satisfied (step S414). The stopping condition may be set such that a preset number of updates is reached, the value of the loss function falls below a first threshold, the number of times the value of the loss function is reduced to or below a second threshold reaches a third threshold, or the like. If it is determined that the stopping condition is not satisfied (step S414: NO), steps S401 to S414 are repeated for new target data x and non-interest data b. By repeating steps S401 to S414, it becomes possible to train the first model parameters and / or the second model parameters so as to reduce the value of the loss function loss, which includes similarities S1, S2, and S3.
[0057] If it is determined in step S414 that the stopping condition is satisfied (step S414: YES), the learning control unit 18 outputs the first model parameters and / or the second model parameters (step S415). The output first model parameters and / or the second model parameters are stored in the storage device 2.
[0058] When step S415 is performed, the expression learning process according to this embodiment ends.
[0059] The processing procedure of the above expression learning processing is one example, and processing can be added, deleted, and / or modified without departing from the gist of the invention.
[0060] As an example, the order of the steps illustrated in FIG. 4 can be changed as appropriate. Specifically, the step of acquiring non-interest data (step S405) may be performed before step S404. The step of updating the representative value S'x (step S411) may be performed at any stage from the step of calculating the latent vector Sx (step S402) to the step of determining whether the stopping condition is satisfied (step S414). Similarly, the step of updating the representative value Z'x (step S412) may be performed at any stage from the step of calculating the non-interest latent vector Zx (step S402) to the step of determining whether the stopping condition is satisfied (step S414), and the step of updating the representative value Z'b (step S413) may be performed at any stage from the step of calculating the non-interest latent vector Zb (step S406) to the step of determining whether the stopping condition is satisfied (step S414). The calculation process of similarity S1 (step S404) may be performed at any stage from the calculation process of latent vector Sx (step S402) and the calculation process of non-interest latent vector Zx (step S403) to the calculation process of loss function (step S410). Similarly, the calculation process of similarity S2 (step S407) may be performed at any stage from the calculation process of non-interest latent vector Zb (step S407) to the calculation process of loss function (step S410), and the calculation process of similarity S3 (step S408) may be performed at any stage from the calculation process of non-interest latent vector Zx (step S403) to the calculation process of loss function (step S410).
[0061] As another example, the loss function does not need to include all of similarity S1, similarity S2, and similarity S3. For example, the loss function may include similarity S1 and similarity S2 but not similarity S3. As another example, the loss function may further include the mutual information between the latent vector Sx and the first non-interest latent vector Zx.
[0062] According to the above-described representation learning process, the first model parameters and / or the second model parameters are updated based on a loss function including similarity S1 and similarity S2. Similarity S1 is an index obtained by correcting the similarity between a latent vector Sx and its representative value S'x using the similarity between a non-interest latent vector Zx and its representative value Z'x. Similarity S2 is an index representing the similarity between a non-interest latent vector Zb and its representative value Z'b. By using such a loss function, it becomes possible to perform representation learning for target data having a feature of interest and a feature of non-interest, emphasizing the feature of interest and suppressing the feature of non-interest.
[0063] An example of how to use the information resource obtained by the expression learning process will be described below.
[0064] <Usage example 1> In the usage example 1, clustering is performed based on the feature of interest of the target data x used in the expression learning process according to this embodiment. The post-processing unit 19 according to the usage example 1 clusters the target data x based on the set of latent vectors Sx calculated by the expression learning process and stored in the Sx memory 21.
[0065] 8 is a diagram showing the processing procedure of clustering according to Use Example 1. As shown in FIG. 8, the post-processing unit 19 acquires P latent vectors Sx from the Sx memory 21 (step S801). P is a natural number equal to or greater than 2. The P latent vectors Sx stored in the Sx memory 21 have been calculated using the first model parameters in the process of the expression learning process.
[0066] After step S801, post-processing unit 19 clusters P pieces of target data x using latent vector Sx (step S802). Clustering is performed by unsupervised clustering. Specifically, clustering is performed using the Kmeans algorithm. Specifically, post-processing unit 19 initially assigns one of multiple labels to each of P latent vectors Sx in the latent space (step A). For each label, post-processing unit 19 calculates the centroid of multiple latent vectors Sx belonging to that label (step B). For each of P latent vectors Sx, post-processing unit 19 calculates the distance from the multiple centroids, selects the label to which the centroid corresponding to the shortest distance belongs, and re-assigns the selected label to that latent vector Sx (step C). Post-processing unit 19 repeats steps A to C until the label assigned in step C does not change. In this way, the latent vectors Sx are clustered. Since there is a one-to-one correspondence between the latent vector Sx and the target data x, clustering of the latent vector Sx also results in clustering of the target data x.
[0067] Here, we will explain the differences between clustering according to this embodiment and clustering according to a comparative example. In clustering, classification is performed using the distance between latent vectors or the distance between a latent vector and a cluster center. In the expression learning process according to the comparative example, only the first vector calculation unit 13 according to this embodiment is used, and the second vector calculation unit 14 is not used. As a result, the machine learning model learns both target features that are of interest and non-target features that are not of interest indistinguishably. For example, when clustering first data and second data that have the same target features but different non-target features, the first data and second data will be classified into different classes.
[0068] The expression learning process according to this embodiment uses both the first vector calculation unit 13 and the second vector calculation unit 14. This makes it possible to obtain a first machine learning model that extracts only the feature of interest and a second machine learning model that extracts only the non-features of interest that are not desired to be focused on. Therefore, even when clustering first data and second data that have the same feature of interest but different non-features of interest, it is possible to classify the first data and the second data into the same class by focusing only on the feature of interest.
[0069] FIG. 9 shows the clustering accuracy of this embodiment and a comparative example. The comparative example in FIG. 9 is the technique described in Non-Patent Document 1. FIG. 9 shows the classification accuracy of the publicly available handwritten digit image dataset Mnist and the classification accuracy of the background-attached Mnist dataset, which adds vertical, horizontal, and diagonal stripes to Mnist. Mnist is data of handwritten digits from 0 to 9, consisting of 10 classes. Mnist with background adds four types of stripes (horizontal lines, vertical lines, 45-degree diagonal lines, and -45-degree diagonal lines) to Mnist, totaling 40 classes. As for clustering accuracy, the classification performance of 10 classes is shown, focusing on numbers. While the accuracy of Mnist in the comparative example is 97.8%, the accuracy of Mnist with background is 34.0%, which is significantly degraded. On the other hand, the classification performance of this embodiment is significantly improved to 97.2%.
[0070] Figure 10 shows the results of visualizing the data distribution in the latent space compressed into two dimensions by this embodiment and a comparative example. The comparative example in Figure 10 is also based on the technique described in Non-Patent Document 1. Each point represents an image of the target data, the hatching in the upper figure represents the number label, and the hatching in the lower figure represents the background label. In the comparative example, data is clustered according to the number label, but data is also clustered according to the background label, and even the same number is distributed separately. In the latent space trained in this embodiment, while data is clustered according to the number label, data is distributed regardless of the background label. Therefore, a latent space that does not include non-interest features that we want to suppress is learned, and we can expect to improve classification accuracy.
[0071] <Usage example 2> In Usage Example 2, clustering is performed using the feature of interest of new data x' that has not been used in the expression learning process according to this embodiment. The post-processing unit 19 according to Usage Example 2 calculates a latent vector Sx' in the latent space of the new data x' using the first model parameters, and clusters the new data x' based on the calculated latent vector Sx'.
[0072] FIG. 11 is a diagram showing the processing procedure of clustering according to Usage Example 2. As shown in FIG. 11, the post-processing unit 19 acquires P pieces of new data x' (step S1101). After step S1101 is performed, the post-processing unit 19 calculates a latent vector Sx' of the new data x' using the first model parameters (step S1102). After step S1102 is performed, the latent vector Sx' is used to cluster the P pieces of new data x' (step S1103). The clustering may be performed using the same method as in step S802.
[0073] According to the second use example, clustering can be performed on new data x' that has not been used in the expression learning process using the first model parameters trained in the expression learning process. Therefore, clustering with higher accuracy can be performed compared to the comparative example.
[0074] <Usage example 3> In the use example 3, clustering is performed using non-interest features of the target data x used in the expression learning process according to this embodiment. The post-processing unit 19 according to the use example 3 clusters the target data x based on a set of non-interest latent vectors Zx.
[0075] FIG. 12 is a diagram showing the processing procedure of clustering according to Use Example 3. As shown in FIG. 12, the post-processing unit 19 acquires P non-interest latent vectors Zx from the Zx memory 22 (step S1201). The P latent vectors Zx stored in the Zx memory 22 are calculated using the second model parameters in the process of the above-mentioned representation learning process. After step S1201 is performed, the post-processing unit 19 clusters the P target data x using the non-interest latent vectors Zx (step S1202). The clustering may be performed in the same manner as step S802, with the "latent vector Sx" replaced with the "non-interest latent vector Zx."
[0076] According to the use example 3, clustering can be performed using the target data x and non-interest latent vector Zx used in the expression learning process. Therefore, clustering with higher accuracy can be performed compared to the comparative example.
[0077] <Usage example 4> In the fourth application example, clustering is performed on new data x' using non-interest features. The post-processing unit 19 according to the fourth application example calculates a non-interest latent vector Zx' in the latent space of the new data x' using the second model parameters, and clusters the new data x' based on the calculated non-interest latent vector Zx'.
[0078] FIG. 13 is a diagram showing the processing procedure of clustering according to Use Example 4. As shown in FIG. 13, the post-processing unit 19 acquires P pieces of new data x′ (step S1301). After step S1301 is performed, the post-processing unit 19 calculates a non-interest latent vector Zx′ of the new data x′ using the second model parameters (step S1302). After step S1302 is performed, the non-interest latent vector Zx′ is used to cluster the P pieces of new data x′ (step S1303). The clustering may be performed in the same manner as step S802, with the “latent vector Sx” replaced with the “non-interest latent vector Zx′”.
[0079] According to Use Example 4, clustering can be performed on new data x' that has not been used in the expression learning process using the second model parameters trained in the expression learning process. Therefore, clustering with higher accuracy than in the comparative example can be performed.
[0080] <Usage example 5> In the fifth example, target data similar to new data is searched for based on a feature of interest. The post-processing unit 19 in the fifth example calculates a new latent vector Sx' in the latent space of the new data x' using the first model parameters, and calculates the distance or similarity between the latent vector Sx' and the latent vector Sx. The display control unit 20 displays the target data x on the display device 5 in order of decreasing distance or similarity.
[0081] FIG. 14 is a diagram showing the processing steps of the search process according to Use Example 5. As shown in FIG. 14, the post-processing unit 19 acquires P new data x' (step S1401). After step S1401, the post-processing unit 19 calculates a latent vector Sx' of the new data x' using the first model parameters (step S1402). After step S1402, the post-processing unit 19 acquires P latent vectors Sx from the Sx memory 21 (step S1403). After step S1403, the post-processing unit 19 calculates the distance between the latent vector Sx' and the latent vector Sx (step S1404). The distance means the difference between the latent vector Sx' and the latent vector Sx in the latent space.
[0082] When step S1404 is performed, the display control unit 20 presents target data x similar to the new data x' from among the P target data x (step S1405). In step S1405, as an example, the display control unit 20 displays on the display device 5 the target data x whose distance is equal to or less than a threshold as target data x similar to the new data x'. At this time, the display control unit 20 may display the target data x similar to the new data x' in a ranking format in order of shortest distance. As another example, the display control unit 20 may display all or a predetermined number of the target data x in order of shortest distance. Note that in steps S1404 to S1405, the similarity, cosine similarity, or other similarity used in the above expression learning process may be used instead of distance.
[0083] According to the use example 5, the first model parameters and latent vector Sx obtained by the expression learning process can be used to search for target data x similar to new data x'. Therefore, the accuracy of the search process is expected to be improved.
[0084] <Usage example 6> In Usage Example 6, other new data x' similar to new data x is searched for based on a feature of interest. The post-processing unit 19 in Usage Example 6 uses the first model parameters to calculate a first new latent vector Sx1' in the latent space of the first new data x1' and multiple second new latent vectors Sx2' in the latent space of multiple second new data x2', and calculates the distance or similarity between each of the multiple second new latent vectors Sx2' and the first new latent vector Sx1'. The display control unit 20 displays the multiple second new data x2' on the display device 5 in order of decreasing distance or similarity.
[0085] FIG. 15 is a diagram showing the processing procedure of the search process according to the sixth use example. As shown in FIG. 15, the post-processing unit 19 acquires a first new data item x1′ and P second new data items x2′ (step S1501). After step S1501, the post-processing unit 19 calculates a latent vector Sx1′ of the new data item x1′ and a latent vector Sx2′ of the new data item x2′ using the first model parameters (step S1502). After step S1502, the post-processing unit 19 calculates the distance between the latent vector Sx1′ and the latent vector Sx2′ (step S1503). After step S1503, the display control unit 20 presents new data item x2′ similar to the new data item x1′ from among the P new data items x2′ (step S1504). The method of presenting the new data item x2′ similar to the new data item x1′ may be, as in step S1405, a ranking format or the like. As in the fifth use example, similarity may be used instead of distance.
[0086] According to Use Example 6, new data x2' similar to new data x1' can be searched for using the first model parameters obtained by the expression learning process. Therefore, improvement in the accuracy of the search process is expected.
[0087] <Usage example 7> In the seventh application example, target data similar to new data is searched for based on non-interest features. The post-processing unit 19 according to the seventh application example calculates multiple new non-interest latent vectors Zx' in the latent space of multiple new data x' using the second model parameters, calculates the distance or similarity between each of the multiple new non-interest latent vectors Zx' and the non-interest latent vector Zx of the target data x, and the display control unit 20 displays the multiple new data on the display device in order of decreasing distance or similarity.
[0088] FIG. 16 is a diagram showing the processing procedure of the search process according to the seventh example of use. As shown in FIG. 16, the post-processing unit 19 acquires new data x′ (step S1601). After step S1601, the post-processing unit 19 calculates a latent vector Sx′ of the new data x′ using the model parameters (step S1602). After step S1602, the post-processing unit 19 acquires P latent vectors Zx of no interest from the Zx memory 22 (step S1403). After step S1603, the post-processing unit 19 calculates the distance between the latent vector Sx′ and the latent vector Zx of no interest (step S1604). After step S1604, the display control unit 20 presents target data x similar to the new data x′ from among the P target data x (step S1605). The target data x similar to the new data x′ may be presented in a ranking format, as in step S1405. As in the fifth example of use, similarity may be calculated instead of distance.
[0089] According to the use example 6, the second model parameters and the non-interest latent vector Zx obtained by the representation learning process can be used to search for target data x similar to the new data x'. Therefore, the accuracy of the search process is expected to be improved.
[0090] <Usage example 8> In Usage Example 8, other new data similar to new data is searched for based on non-interest features. The post-processing unit 19 in Usage Example 8 uses the second model parameters to calculate a first new non-interest latent vector Zx1′ in the latent space of the non-interest features of the first new data x1′ and multiple second new non-interest latent vectors Zx2′ in the latent space of the non-interest features of the multiple second new data x2′, calculates the distance or similarity between each of the multiple second new non-interest latent vectors Zx2′ and the first new non-interest latent vector Zx1′, and the display control unit 20 displays the multiple second new data x2′ on the display device 5 in order of closest distance or similarity.
[0091] FIG. 17 is a diagram showing the processing procedure of the search process according to the eighth use example. As shown in FIG. 17, the post-processing unit 19 acquires the first new data x1′ and P pieces of second new data x2′ (step S1701). After step S1701 is performed, the post-processing unit 19 calculates the non-focus latent vector Zx1′ of the new data x1′ and the non-focus latent vector Zx2′ of the new data x2′ using the second model parameters (step S1702). After step S1702 is performed, the post-processing unit 19 calculates the distance between the non-focus latent vector Zx1′ and the non-focus latent vector Zx2′ (step S1703). After step S1703 is performed, the display control unit 20 presents new data x2′ similar to the new data x1′ from among the P pieces of new data x2′ (step S1704). The method of presenting the new data x2′ similar to the new data x1′ may be, as in step S1405, a ranking format or the like. Similar to the use case 5, similarity may be calculated instead of distance.
[0092] According to Use Example 8, new data x2' similar to new data x1' can be searched for using the second model parameters obtained by the expression learning process. Therefore, improvement in the accuracy of the search process is expected.
[0093] (Summary) As in the various embodiments described above, the expression learning device 100 includes a first acquisition unit 11, a second acquisition unit 12, a first vector calculation unit 13, a second vector calculation unit 14, a similarity calculation unit 15, a loss function calculation unit 16, and an update unit 17. The first acquisition unit 11 acquires target data x. The second acquisition unit 12 acquires non-target data b similar to a non-target feature included in the target data x. The first vector calculation unit 13 calculates a latent vector Sx in the latent space of the target data x using first model parameters related to a first machine learning model to be trained. The second vector calculation unit 14 calculates a non-target latent vector Zx in the latent space of the non-target feature included in the target data x and a non-target latent vector Zb in the latent space of the non-target data b using second model parameters related to a second machine learning model to be trained. The similarity calculation unit 15 calculates a similarity S1 by correcting the similarity between the latent vector Sx and its representative value S'x with the similarity between the non-interest latent vector Zx and its representative value Z'x, and calculates a similarity S2 between the non-interest latent vector Zb and its representative value Z'b. The loss function calculation unit 16 calculates a loss function including the similarity S1 and the similarity S2. The update unit 17 updates the first model parameter and the second model parameter based on the loss function.
[0094] According to the above configuration, the first model parameters and / or the second model parameters are updated based on a loss function including similarity S1 and similarity S2. Similarity S1 is an index obtained by correcting the similarity between a latent vector Sx and its representative value S'x using the similarity between a non-interest latent vector Zx and its representative value Z'x. Similarity S2 is an index representing the similarity between a non-interest latent vector Zb and its representative value Z'b. By using such a loss function, it becomes possible to perform representation learning for target data having a feature of interest and a feature of non-interest, emphasizing the feature of interest and suppressing the feature of non-interest.
[0095] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]
[0096] 1...processing circuit, 2...memory device, 3...input device, 4...communication device, 5...display device, 11...first acquisition unit, 12...second acquisition unit, 13...first vector calculation unit, 14...second vector calculation unit, 15...similarity calculation unit, 16...loss function calculation unit, 17...update unit, 18...learning control unit, 19...post-processing unit, 20...display control unit, 100...representation learning device.
Claims
1. a first acquisition unit that acquires target data; a second acquisition unit that acquires non-interest data similar to the non-interest feature included in the target data; a first vector calculation unit that calculates a latent vector in a latent space of the target data using first model parameters related to a first machine learning model to be trained; a second vector calculation unit that calculates a first non-interest latent vector in a latent space of a non-interest feature included in the target data and a second non-interest latent vector in a latent space of the non-interest data using second model parameters related to a second machine learning model to be trained; a similarity calculation unit that calculates a first similarity obtained by correcting the similarity between the latent vector and a first representative value of the latent vector by the similarity between the first non-interest latent vector and a second representative value of the first non-interest latent vector, and a second similarity between the second non-interest latent vector and a third representative value of the second non-interest latent vector; a loss function calculation unit that calculates a loss function including the first similarity and the second similarity; an update unit that updates the first model parameters and / or the second model parameters based on the loss function; An expression learning device comprising:
2. the similarity calculation unit calculates a third similarity between the first non-interest latent vector and the third representative value; the loss function calculation unit calculates the loss function including the third similarity in addition to the first similarity and the second similarity.
2. The expression learning device according to claim 1.
3. The expression learning device according to claim 1 , further comprising a first storage unit that stores the first representative value.
4. The expression learning device according to claim 3 , wherein the first representative value is a moving average of the set of latent vectors calculated by the first vector calculation unit.
5. The expression learning device according to claim 1 , further comprising a second storage unit that stores the second representative value.
6. The expression learning device according to claim 5 , wherein the second representative value is a moving average of the set of the first non-interest latent vectors calculated by the second vector calculation unit.
7. 3. The expression learning device according to claim 2, further comprising a third storage unit that stores the third representative value.
8. The expression learning device according to claim 7 , wherein the third representative value is a moving average of the set of the second non-interest latent vectors calculated by the second vector calculation unit.
9. 3. The representation learning device according to claim 2, wherein the loss function calculation unit calculates the loss function further including a first feature decorrelation term representing a degree to which the latent vectors are orthogonal to one another, a second feature decorrelation term representing a degree to which the first non-interest latent vectors are orthogonal to one another, and a third feature decorrelation term representing a degree to which the second non-interest latent vectors are orthogonal to one another.
10. The representation learning device according to claim 2 , wherein the loss function calculation unit calculates the loss function further including mutual information between the latent vector and the first non-interest latent vector.
11. The representation learning device according to claim 1 , further comprising a post-processing unit that clusters the target data based on the set of latent vectors.
12. The representation learning device according to claim 1 , further comprising a post-processing unit that calculates new latent vectors in a latent space of new data using the first model parameters, and clusters the new data based on the new latent vectors.
13. The representation learning device according to claim 1 , further comprising a post-processing unit that clusters the target data based on the first set of non-interest latent vectors.
14. The representation learning device according to claim 1 , further comprising a post-processing unit that calculates new non-interest latent vectors in a latent space of new data using the second model parameters, and clusters the new data based on the new non-interest latent vectors.
15. a post-processing unit that calculates a new latent vector in a latent space of new data using the first model parameters and calculates a distance or a similarity between the new latent vector and the latent vector; a display control unit that displays the target data on a display device in order of decreasing distance or similarity; 2. The expression learning device according to claim 1.
16. a post-processing unit that uses the first model parameters to calculate a first new latent vector in a latent space of the first new data and a plurality of second new latent vectors in a latent space of a plurality of second new data, and calculates a distance or a similarity between each of the plurality of second new latent vectors and the first new latent vector; a display control unit that displays the plurality of second new data on a display device in order of decreasing distance or similarity; 2. The expression learning device according to claim 1.
17. a post-processing unit that calculates a plurality of new first non-interest latent vectors in a latent space of a plurality of new data using the second model parameters, and calculates a distance or a similarity between each of the plurality of new first non-interest latent vectors and the first non-interest latent vector; a display control unit that displays the plurality of new data on a display device in order of decreasing distance or similarity, 2. The expression learning device according to claim 1.
18. a post-processing unit that uses the second model parameters to calculate a first new non-interest latent vector in a latent space of a non-interest feature of the first new data and a plurality of second new non-interest latent vectors in a latent space of a non-interest feature of the plurality of second new data, and calculates a distance or a similarity between each of the plurality of second new non-interest latent vectors and the first new non-interest latent vector; a display control unit that displays the plurality of second new data on a display device in order of decreasing distance or similarity; 2. The expression learning device according to claim 1.
19. Obtain the target data, acquiring non-interest data similar to the non-interest features included in the target data; Calculating latent vectors in a latent space of the target data using first model parameters for a first machine learning model to be trained; calculating a first latent vector of non-interest in a latent space of non-interest features included in the target data and a second latent vector of non-interest in a latent space of the non-interest data using second model parameters related to a second machine learning model to be trained; calculating a first similarity obtained by correcting the similarity between the latent vector and a first representative value of the latent vector by the similarity between the first non-interest latent vector and a second representative value of the first non-interest latent vector, and a second similarity between the second non-interest latent vector and a third representative value of the second non-interest latent vector; calculating a loss function including the first similarity and the second similarity; updating the first model parameters and / or the second model parameters based on the loss function; A computer-implemented representation learning method comprising:
20. On the computer, A function to acquire target data; a function of acquiring non-interest data similar to the non-interest features included in the target data; a function of calculating a latent vector in a latent space of the target data using first model parameters related to a first machine learning model to be trained; a function of calculating a first non-interest latent vector in a latent space of a non-interest feature included in the target data and a second non-interest latent vector in a latent space of the non-interest data using second model parameters related to a second machine learning model to be trained; a function of calculating a first similarity obtained by correcting the similarity between the latent vector and a first representative value of the latent vector by the similarity between the first non-interest latent vector and a second representative value of the first non-interest latent vector, and a second similarity between the second non-interest latent vector and a third representative value of the second non-interest latent vector; a function of calculating a loss function including the first similarity and the second similarity; a function of updating the first model parameters and / or the second model parameters based on the loss function; An expression learning program that makes this possible.
Citation Information
Patent Citations
CLR2021
Learning device, estimation device, data generation device, learning method and learning program
JP2020149504A
Universal feature representation learning for face recognition
US20210142043A1
Disentangled Representations For Gait Recognition
US20220148335A1
Domain adaptation and fusion using task-irrelevant paired data in sequential form
WO2020256732A1
Cited By
Defect classification support device, method, and program
JP2025034191A
Defect classification support device, method, and program
JP7892634B2