Data storage system

The data storage system addresses the inefficiency in storing compressed data by using an outlier detection and learning unit to focus on main data for training a data generation model, resulting in reduced storage needs and improved data reconstruction accuracy.

JP2025074930APending Publication Date: 2025-05-14DENSO CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024105628
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-06-28
Publication Date
2025-05-14

AI Technical Summary

Technical Problem

In data storage systems that use deep generation models like VAE and GAN for data compression, the amount of compressed data increases with the input data, leading to inefficient storage capacity and reduced accuracy in reconstructing input data.

Method used

A data storage system that includes an outlier detection unit to separate main data from outliers, a learning unit to train a data generation model using only the main data, and a data storage unit that stores the learned data generation model and detected outliers, allowing for efficient storage and improved data reconstruction accuracy.

Benefits of technology

This approach reduces the amount of data stored by eliminating the need for compressed data and improves the accuracy of data reconstruction by focusing on frequently generated main data and storing outliers separately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025074930000001_ABST
    Figure 2025074930000001_ABST
Patent Text Reader

Abstract

To provide a data storage system configured to construct a data generative model that can properly restore input data without storing compressed input data.SOLUTION: A data storage system includes an outlier detection unit 12, a training unit 14, and a data storage unit 16. The outlier detection unit detects an outlier which is isolated from a main data group and appears with low frequency, from input data. The training unit trains a data generative model by inputting main data obtained by excluding the outlier from the input data, to the data generative model. The data storage unit stores the trained data generative model and the outlier detected by the outlier detection unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a data storage system. [Background technology]

[0002] As described in Patent Document 1, a transfer system is known that includes a data compression unit constructed using a deep generative model such as VAE or GAN, and is configured to compress input data into compressed data in the form of a multidimensional joint probability distribution in the data compression unit and transmit the compressed data to a data reproduction unit.

[0003] VAE stands for "Variational Auto Encoder", and GAN stands for "Generative Adversarial Net". [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2018-61091 A Summary of the Invention [Problem to be solved by the invention]

[0005] In the above-mentioned transfer system, the amount of compressed data increases as the amount of input data increases, so there was a problem that the amount of compressed data sent to the data reproduction unit could not be reduced, and the storage capacity for storing compressed data in the data reproduction unit increased.

[0006] An object of one aspect of the present disclosure is to provide a data storage system capable of constructing a data generation model capable of properly restoring input data without storing compressed data obtained by compressing the input data. [Means for solving the problem]

[0007] A data storage system according to one embodiment of the present disclosure includes an outlier detection unit (12), a learning unit (14), and a data storage unit (42). Here, the outlier detection unit detects outliers from the input data that are isolated from the main data group and occur infrequently. The learning unit inputs the main data, excluding the outliers from the input data, to a data generation model using a machine learning model to learn the data generation model. The data storage unit stores the data generation model learned by the learning unit and the outliers detected by the outlier detection unit.

[0008] Typically, when training a data generation model, all of the data obtained is used for training. However, because outlier data often has different characteristics from normal data, it is generally difficult to reproduce similar data even after training, resulting in a decrease in accuracy during generation. Furthermore, as data generation models attempt to train in a way that allows them to generate data that includes outlier data, the accuracy of data other than outliers also tends to decrease after generation. In addition, the main data is frequently occurring data, and storing all of this data is considered inefficient from the perspective of data variety and data volume.

[0009] Therefore, in the technology disclosed herein, the learning unit creates a data generation model that can reconstruct the input data by learning using main data that occurs frequently, with outliers removed from the input data, and the reconstruction accuracy of the input data by the data generation model is improved. Also, since outliers detected by the outlier detection unit are generally likely to have poor reconstruction accuracy, the data storage unit stores the outliers as originals in addition to the data generation model.

[0010] Therefore, according to the data storage system of the present disclosure, it becomes possible to restore input data using a stored data generation model without transferring or storing compressed data obtained by compressing input data as described in Patent Document 1. In addition, by adding outliers stored by the data storage unit to the restored input data, it is possible to improve the accuracy of restoring the input data.

[0011] As a technology for transmitting and storing a portion of input data as is, such as the data storage system of the present disclosure, there are known technologies described in reference documents such as WO2022 / 074794A1 and WO2021 / 250868A1. However, the data storage system of the present disclosure is completely different from the technologies described in the reference documents in that it stores a data generation model and outliers.

[0012] In other words, the technology described in the reference document is a technology in which, in a system in which input data is compressed using an autoencoder's encoder and then transmitted, input data that cannot be properly restored by the autoencoder's decoder is transmitted as is, without being compressed.

[0013] In this technology, input data that cannot be restored correctly by the encoder is sent as is because the input data is unlearned input data that has not been used in learning the autoencoder. In other words, if unlearned input data is compressed and sent, the input data cannot be restored on the receiving side, so the unlearned input data is sent without compression.

[0014] Therefore, in the technology described in the above reference, the data stored on the receiving side is compressed data and unlearned input data, which is similar to the technology disclosed herein in that the unlearned input data is stored as is.

[0015] However, in the data storage system of the present disclosure, input data and compressed data other than the outliers are not transferred or stored, and the data generation model trained using the main data is stored.

[0016] Therefore, according to the data storage system disclosed herein, the amount of data stored by the data storage unit can be reduced, making the storage capacity extremely small, compared to those described in Patent Document 1 and the above-mentioned references. [Brief description of the drawings]

[0017] [Figure 1] 1 is a block diagram showing a configuration of a data storage system according to a first embodiment. [Diagram 2] 1 is an explanatory diagram showing the learning operation of a VAE in a learning unit and a data generation model obtained by the learning. [Diagram 3] 4 is an explanatory diagram illustrating an outlier detected by an outlier detection unit. FIG. [Figure 4] FIG. 11 is a block diagram showing the configuration of a data storage system according to a second embodiment. [Diagram 5] 13 is a flowchart showing a data storage process executed by a computer in the third embodiment. [Figure 6] 13 is a flowchart showing a data storage process executed by a server in the third embodiment. [Figure 7] FIG. 13 is a block diagram showing the configuration of a data storage system according to a fourth embodiment. [Figure 8] FIG. 13 is a block diagram showing the configuration of a data storage system according to a fifth embodiment. [Figure 9] 13 is a flowchart showing an outlier detection process executed in a learning unit of the fifth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. [First embodiment] As shown in FIG. 1, a data storage system 1 of this embodiment includes an input unit 2, an outlier detection unit 12, a learning unit 14, a data storage unit 16, and a memory unit 20.

[0019] The input unit 2 is for inputting a large amount of input data, such as image data, and an outlier ratio (to be described later), from the outside, to a computer 10 including a CPU, a ROM, and a RAM.

[0020] The outlier detection unit 12, the learning unit 14, and the data storage unit 16 are functions of the computer 10 that are realized by the CPU of the computer 10 executing a predetermined program.

[0021] The storage unit 20 is composed of a storage medium such as a memory, a hard disk, or an SSD, or a server for storing data capable of communicating with the computer 10, and is capable of storing various data described below.

[0022] Here, the outlier detection unit 12 detects outliers that are isolated from the main data group and occur infrequently from the input data, and separates the input data into outlier data and main data. This separation is performed based on a preset outlier ratio based on an input from the input unit 2. In other words, the outlier detection unit 12 detects outliers and separates them from the input data so that the ratio between the number of outliers and the number of main data becomes the preset outlier ratio.

[0023] Then, the outlier data separated from the input data by the outlier detection unit 12 is input to the data storage unit 16, and the main data is input to the learning unit . The learning unit 14 inputs the main data into a data generation model using a machine learning model to learn the data generation model. Specifically, in this embodiment, the learning unit 14 uses the above-mentioned VAE as the data generation model.

[0024] 2, the VAE is a learning device composed of an encoder 32 and a decoder 34, which are usually composed of a neural network. If the encoding result of the input x is z, then the encoding result z is converted to a decoding result x' by the decoder 34. Then, the VAE learns the internal parameters of the encoder 32 and the decoder 34 so that the input x becomes the decoding result x'.

[0025] Below, we will explain the learning of VAE. In the following explanation, the input data is (x1, , xn), and the intermediate layer data corresponding to each input data is (z1, , zn). Furthermore, capital letters such as X and Z represent random variables.

[0026] Furthermore, pe(x) represents the empirical distribution, and pβ(x) represents the generalized marginal likelihood, and their definitions are as follows:

[0027]

number

[0028] Furthermore, p(x|z) and q(z|x) are probability distributions obtained by the decoder 34 and the encoder 32, respectively. For example, in this embodiment, p(x|z) is given by the following equation.

[0029]

number

[0030] Here, D represents the dimension of the data. Also, q(z|x) is a normal distribution determined for each data used in VEA. p(z) is the probability distribution of the target intermediate layer, and although the standard normal distribution can be used, other distributions, such as a mixed normal distribution, can also be used.

[0031] Also,

[0032]

number

[0033]

number

[0034] Next, VAE learning is achieved by minimizing the following equation (1).

[0035]

number

[0036] In equation (1), the first term is the negative log-likelihood and acts to reduce the reconstruction error. On the other hand, the second term acts as a regularization term with the aim of bringing the intermediate distribution closer to a known distribution p(z). For example, a standard normal distribution can be used as p(z). Then, by minimizing equation (1), an encoder 32 and a decoder 34 are trained such that the intermediate distribution (distribution of z) 36 is known (here, normal distribution).

[0037] 2, the learning unit 14 outputs the decoder portion 34 and normal distribution 38 learned as described above to the data storage unit 16 as the data generation model 30. The learning unit 14 also outputs to the data storage unit 16 the number of main data used to learn this data generation model 30.

[0038] According to the data generation model 30 thus output from the learning unit 14 to the data storage unit 16, random numbers are generated from a normal distribution 38, which is a data distribution, and input to the decoder 34, so that data x' can be output from the decoder 34. In other words, a data distribution similar to the data distribution constituting the input data x can be restored.

[0039] The learning of the VAE in the learning unit 14 is described in the above-mentioned references and is a publicly known technology, so detailed description will be omitted here. In addition, in this embodiment, the VAE is taken up as the data generation model 30, but for example, the above-mentioned GAN and its improvement method, a generation model using a diffusion model (Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models, 2020.), etc. can also be used.

[0040] Next, outlier detection in the outlier detection unit 12 is realized by minimizing the following equation (2).

[0041]

number

[0042] In equation (2), the first and second terms on the right-hand side differ from equation (1) in that β and πi are added as coefficients, but their roles are the same: the first term acts to reduce the reconstruction error, and the second term acts as a regularization term to bring it closer to the known distribution. β is a constant for balancing the reconstruction error and regularization.

[0043] Furthermore, the third and fourth terms are necessary to form an upper bound on the left side of the equation, but in this embodiment, these terms are ignored since they have little effect on the minimization of equation (2). Therefore, the only difference between equation (2) and equation (1) is the balance coefficient β and the presence of weights πi for each data item in the first and second terms.

[0044] The right-hand side is a quantity determined for two probability distributions p(x) and q(x) used in robust statistics, called gamma entropy, and is given by the following equation (3).

[0045]

number

[0046] In this embodiment, the fact that equation (3) is a value equivalent to the robust distance between the two distributions is utilized, and by minimizing the right-hand side of equation (2), the left-hand side also becomes small. Moreover, πi is calculated by the following equation (4).

[0047]

number

[0048] The outlier detection unit 12 then calculates πi in equation (4) for each data xi in addition to minimizing equation (2). That is, the outlier detection unit 12 alternately repeats the minimization of equation (2) and the calculation of πi using equation (4), and detects outliers using πi obtained as a result of the calculation as a parameter representing the normality of each data xi.

[0049] For example, Figure 3 shows a histogram of normality πi obtained by performing the above calculation process on input data x, which is a mixture of images of “0” from a database of handwritten digits called MNIST, and 1% of images of “1”.

[0050] As is clear from Figure 3, the gray input data x, which represents "0," is located in an area with a high normality πi, while the black input data x, which represents "1," is distributed in an area with a normality of 0. In addition, gray input data x is also distributed in part of the area with a normality of 0, but this is an image that is different from a normal "0," such as a crushed "0."

[0051] For this reason, in the outlier detection unit 12, based on a preset threshold value for the normality πi calculated for each input data xi, input data whose normality πi is less than the threshold value is detected as outlier data and separated from the main data.

[0052] In this embodiment, the threshold value for detecting outliers is appropriately changed so that the ratio of the main data to the outlier data corresponds to the outlier ratio input from the input unit 2. Next, the data storage unit 16 stores the data generation model 30 input from the learning unit 14 and the outlier data 22 input from the outlier detection unit 12 in a predetermined storage area of ​​the storage unit 20.

[0053] Furthermore, the data storage unit 16 stores the number of main data input from the learning unit 14 and the number of outlier data 22 in a predetermined storage area of ​​the memory unit 20 as the number of data 24 used to generate data in the data generation model 30. Note that the data storage unit 16 may be configured to store the number of input data in the memory unit 20 instead of the number of main data.

[0054] As described above, according to the data storage system 1 of this embodiment, the data generation model 30 generated by the learning unit 14, the outlier data 22, and the number of data 24 of the main data and the outlier data are stored in the memory unit 20.

[0055] As a result, the input data can be restored using the data generation model 30 stored in the memory unit 20, and there is no need to transmit or store compressed data obtained by compressing the input data in order to restore the input data, as described in Patent Document 1 and the above-mentioned references.

[0056] Furthermore, since the outlier data 22 is stored in the storage unit 20, the outlier data 22 can be added to the data restored using the data generation model 30, thereby improving the accuracy of data restoration.

[0057] Furthermore, when restoring input data using the data generation model 30, the number of input data to be restored can be set in accordance with the number of outlier data based on the number of data 24 stored in the memory unit 20, thereby making it possible to more accurately restore the original input data used for learning.

[0058] Furthermore, the storage unit 20 only stores the data generation model 30, the outlier data 22, and the number of pieces of data 24, and does not need to store compressed data of the input data. Therefore, the amount of data stored in the storage unit 20 can be reduced, and the storage capacity of the storage unit 20 can be made smaller.

[0059] Furthermore, since there is no need to transfer compressed data of input data and store it in memory unit 20 as described in Patent Document 1 and the above-mentioned references, it is possible to prevent compressed data including personal information from leaking to the outside and being misused.

[0060] Furthermore, if the number of outlier data stored in the storage unit 20 is increased, the accuracy of data restoration by the data generation model 30 can be improved, but the amount of outlier data stored in the storage unit 20 increases.

[0061] In contrast, in this embodiment, the ratio between the outlier data 22 separated by the outlier detection unit 12 and the main data can be specified by the outlier ratio input from the input unit 2. Therefore, by appropriately setting the outlier ratio, the user can reduce the amount of outlier data stored in the storage unit 20 while ensuring the accuracy of data restoration by the data generation model 30. Here, the appropriate value is, for example, a value determined by the data generation model taking into consideration the desired accuracy and the increase in data volume due to outliers.

[0062] [Second embodiment] As shown in FIG. 4, the data storage system 1 of this embodiment includes a plurality of clients 40 and a server 50 capable of data communication with each of the clients 40.

[0063] The client 40 and the server 50 are each configured as a computer including a CPU, ROM, and RAM, and are connected via a predetermined communication line, for example, a network such as the Internet.

[0064] The multiple clients 40 include the same outlier detection unit 12 and learning unit 14 as in the first embodiment. Then, the outlier data output from the outlier detection unit 12 and the parameters of the data generation model learned by the learning unit 14 are input to a data transmission unit 42, which transmits the outlier data and the parameters of the data generation model to a server 50.

[0065] On the other hand, the server 50 includes a data storage unit 16 and a storage unit 20 similar to those in the first embodiment. The server 50 also includes a data receiving unit 52 that receives transmission data from the multiple clients 40 and outputs the data to the data storage unit 16.

[0066] In the server 50, the data storage unit 16 updates the data generation model 30 stored in the storage unit 20 by using parameters of the data generation model among the transmission data from each client 40 received by the data receiving unit 52. In addition, the data storage unit 16 stores the outlier data transmitted from each client 40 in the storage unit 20.

[0067] Here, the parameters of the data generation model can be updated, for example, by normalizing the parameters of each client so that the overall weighting is 1 based on the amount of data held by each client, and then adding them up.

[0068] In addition, in the above configuration, the parameters are sent from the client only once, but, for example, the parameters of the data generation model updated by the server can be sent again to the client, and the client can update the parameters of the data generation model of each client with its own data, and then send them back to the server. It is also possible to consider an implementation method in which these operations are repeated a predetermined number of times, or until the parameter updates converge.

[0069] That is, the server 50 uses multiple clients 40 to train a data generation model without collecting distributed data in one place. Then, the server 50 obtains the parameters of each data generation model, which is the result of the training, from each client 40 and updates the data generation model 30, thereby realizing so-called federated learning. For more information on federated learning, see, for example, "H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-efficient learning of deep networks from decentralized data. Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), Vol. 54, , 2017."

[0070] In addition, the server 50 not only generates and updates the data generation model 30 by federated learning, but also acquires outlier data from each client 40 and stores it in the storage unit 20. Therefore, according to the data storage system of this embodiment, it is possible to obtain the same effect as in the first embodiment. That is, the input data can be restored using the data generation model 30 stored in the storage unit 20 of the server 50, and further, by adding the outlier data 22 to the restored data, the accuracy of data restoration can be improved.

[0071] In addition, since the memory unit 20 only stores the data generation model 30 and the outlier data 22, the amount of data stored in the memory unit 20 can be reduced, thereby making it possible to reduce the storage capacity of the memory unit 20.

[0072] Furthermore, only the parameters of the data generation model trained on the client 40 side and outlier data are transmitted from the client 40 to the server 50, and there is no need to transmit input data including personal information or compressed data obtained by compressing the input data. This makes it possible to prevent these data from leaking out from the data transmission path from the client 40 to the server 50 or from the server 50 to the outside and being misused. In addition, since actual data is not transmitted from the client to the server, this is an excellent form in terms of privacy protection.

[0073] In addition, the number of main data and the number of outlier data used to train the data generation model on the client 40 side may be transmitted to the server 50, and on the server 50 side, the data storage unit 16 may store these numbers of data in the memory unit 20.

[0074] [Third embodiment] In the first and second embodiments, the data saving unit 16 saves the data generation model 30 and the outlier data 22 learned by the learning unit 14 or each client 40 in the storage unit 20. Therefore, the data generation model and the outlier data saved in the storage unit 20 are updated sequentially during the input period of the input data.

[0075] However, for example, in a system configured to collect data over several months or years, the distribution of the input data may change due to environmental changes during the input data collection period, etc. In this case, the stored data generation model 30 can restore average input data over the entire data collection period, but cannot restore input data for each season or month.

[0076] Therefore, in the data storage system 1 of this embodiment, the trained data generation model 30 and the outlier data 22 are stored for each predetermined set period, such as one month or several weeks. In addition, time information (e.g., a time stamp) indicating the period, in other words, the time, during which the data generation model 30 was trained is added to the stored data.

[0077] Specifically, in the data storage system 1 of the first embodiment, the computer 10 executes the data storage process in the procedure shown in FIG. That is, in S110, input data and an outlier ratio are acquired from the outside via the input unit 2, and in the following S120, a process is performed as the outlier detection unit 12 that detects outliers from the input data based on the acquired outlier ratio. In addition, in the following S130, a process is performed as the learning unit 14 that learns a data generation model based on main data obtained by excluding outliers from the input data.

[0078] Next, in S140, it is determined whether a preset period (for example, one month) has elapsed since the start of the learning process in S130. If the preset period has not elapsed, the process proceeds to S110, and the processes in S110 to S130 are executed again.

[0079] On the other hand, if it is determined in S140 that the set period has elapsed, the process proceeds to S150, and the data generation model learned by the learning process of S130 within the set period and the outlier data detected in S120 are stored in the memory unit 20 together with a timestamp indicating the current learning time.

[0080] Then, in the next step S160, the process waits for the next learning start timing to arrive by determining whether it is currently the next learning start timing for the data generation model, and when the next learning start timing arrives, the process proceeds to S110 and the above series of processes are executed again.

[0081] As a result, the memory unit 20 stores multiple data generation models 30 trained over a specified period of time and outlier data 22 corresponding to each data generation model 30, each with a timestamp.

[0082] Therefore, when a user restores data using the data generation model 30 stored in the memory unit 20, the user can restore data for a desired period by selecting the data generation model 30 trained for the desired period and the outlier data 22.

[0083] In the data storage system 1 of the second embodiment, the server 50 executes the data storage process according to the procedure shown in FIG. 6, thereby achieving the same effects as those described above. That is, in the server 50, in S210, a process as the data receiving unit 52 is executed to obtain the parameters of the data generation model and the outlier data from the client 40.

[0084] Then, in the next S220, the parameters of the data generation model acquired in S210 are used to update the data generation model 30 stored in the memory unit 20, and the outlier data is stored in the memory unit 20, thereby performing processing as the data storage unit 16.

[0085] Next, in S230, it is determined whether a preset period has elapsed since the start of the data generation model update process by S210 and S220, and if the preset period has not elapsed, the process proceeds to S210 and the processes of S210 and S220 are executed again.

[0086] On the other hand, if it is determined in S230 that the set period has elapsed, the process proceeds to S240, where a timestamp indicating the time of this data collection is assigned to the data generation model 30 and the outlier data 22 stored in the memory unit 20 during the set period, and the data are stored as data for that collection period.

[0087] Then, in the next S240, it is determined whether or not it is currently the timing to start the next data collection, and the process waits for the next data collection start timing to arrive. When the next data collection start timing arrives, the process proceeds to S210, and the above series of processes are executed again.

[0088] As a result, the memory unit 20 stores multiple data generation models 30 trained at predetermined intervals and the outlier data 22 corresponding to each data generation model 30, each with a timestamp, thereby achieving the same effect as described above.

[0089] [Fourth embodiment] As shown in FIG. 7, the data storage system 1 of this embodiment has the same basic configuration as that of the first embodiment shown in FIG. 1, and differs from the first embodiment in that it includes a display unit 4 and the computer 10 has the function of a reconstruction accuracy calculation unit 18.

[0090] Of these, the reconstruction accuracy calculation unit 18 calculates the reconstruction accuracy of data by the data generation model 30 trained by the training unit 14. For example, in the VAE shown in Fig. 2 used in training the data generation model 30, the reconstruction accuracy calculation unit 18 calculates the reconstruction accuracy from the difference between the data x input to the encoder 32 during training and the data x' output from the decoder 34. The difference between the data x input to the encoder 32 during training and the data x' output from the decoder 34 can be calculated using statistics such as the average or maximum value of the error.

[0091] Then, the reconstruction accuracy calculation unit 18 outputs the calculated reconstruction accuracy to the display unit 4, thereby displaying it in a predetermined display area of ​​the display unit 4. In addition, the number of detected outliers, that is, the number of pieces of outlier data, is output from the outlier detection unit 12 to the display unit 4, and the display unit 4 displays the number of pieces of data in a predetermined display area.

[0092] As a result, in the data storage system 1 of this embodiment, the reconstruction accuracy of the data generation model 30 learned by the learning unit 14 and the number of outlier data detected by the outlier detection unit 12 are displayed on the display unit 4.

[0093] Therefore, according to the data storage system 1 of this embodiment, the user can set the outlier ratio taking into consideration the reconstruction accuracy of the data generation model 30 displayed on the display unit 4, the number of outlier data, or the storage capacity of the memory unit 20 that stores the outlier data 22.

[0094] For example, when the reconstruction accuracy of the data generation model 30 is low, the outlier ratio can be increased to improve the reconstruction accuracy of the subsequently trained data generation model 30. Also, for example, as the storage capacity of the storage unit 20 increases, the outlier ratio can be decreased to suppress the storage capacity of the outlier data 22 in the storage unit 20.

[0095] [Fifth embodiment] In each of the above embodiments, the outlier detection unit 12 has been described as calculating the normality πi for each piece of input data xi, and detecting input data x for which the normality πi is less than a threshold value as outlier data.

[0096] In contrast, in this embodiment, as shown in FIG. 8, the function of the outlier detector is realized by the outlier detection process of S400 that is repeatedly executed in the learning section 14 together with the learning process of S300.

[0097] That is, in this embodiment, the learning unit 14 executes the learning process of S300 every time input data x is input from the input unit 2. In this learning process, for example, in the data generation model by VAE shown in Fig. 2, the encoder 32 and the decoder 34 are trained so that the output data x' from the decoder 34 approaches the input data x.

[0098] In contrast, in the outlier detection process of S400, each time the learning process of S300 is executed, it is determined whether the latest input data x used in the learning process of S300 is outlier data, using the procedure shown in FIG.

[0099] That is, in the outlier detection process of S400, first, in S410, the input data x and output data x′ used in training the data generation model in S300 are obtained, and the absolute value of the difference (xx′) between them is calculated as a parameter representing the reconstruction accuracy of the input data x.

[0100] Next, in S420, it is determined whether the reconstruction accuracy of the input data x calculated in S410 is less than a preset threshold value. If the difference between the input data x and the output data x' is small and it is determined in S420 that the reconstruction accuracy of the input data x is equal to or greater than the threshold value, the outlier detection process is terminated and the process proceeds to the learning process of S300.

[0101] On the other hand, if the difference between the input data x and the output data x′ is large and it is determined in S420 that the reconstruction accuracy of the input data x is less than the threshold, the process proceeds to S430, and the latest input data x used in the learning process in S300 is determined to be an outlier.

[0102] In S430, the input data x determined to be an outlier is input to the data storage unit 16 as the outlier data 22. As a result, the outlier data 22 is stored in a predetermined storage area of ​​the storage unit 20.

[0103] Furthermore, if it is determined in S430 that the input data x is an outlier, the process proceeds to S440, where the outlier is removed from the input data so that learning processing is performed based on the input data from which the outlier has been removed, in other words, the main data, in S300. Then, once the outlier has been removed from the input data in S440, the outlier detection processing is terminated and the process proceeds to the learning processing of S300.

[0104] As described above, in this embodiment, whether or not the input data x input from the input unit 2 is an outlier is determined by the learning unit 14 based on the reconstruction accuracy of the input data obtained from the learning result of the data generation model. That is, in this embodiment, input data with low reconstruction accuracy below a threshold cannot be properly restored using the data generation model, so it is determined to be an outlier and excluded from the input data used to learn the data generation model.

[0105] Therefore, the data storage system 1 of this embodiment can also obtain the same effect as the data storage system 1 of the first embodiment described above. Moreover, in this embodiment, it is not necessary to extract outliers from all input data used for learning the data generation model, and it is possible to sequentially determine whether or not input data is an outlier during the learning operation in the learning unit 14. Therefore, the time required for extracting outliers can be shortened.

[0106] In this embodiment, the outlier detection process executed by the learning unit 14 has been described as a modified example of the outlier detection unit 12 in the first embodiment. However, the outlier detection process of this embodiment can also be applied to the learning units 14 in the second to fourth embodiments in the same manner as above, and the same effects as above can be obtained.

[0107] Also in the first to fourth embodiments, similarly to this embodiment, by repeatedly performing extraction of outliers and learning of a data generation model, it is possible to construct a data generation model with higher reconstruction accuracy.

[0108] [Other embodiments] Although the embodiments of the present disclosure have been described above, the present disclosure is not limited to the above-described embodiments and can be implemented in various modified forms.

[0109] For example, in the first embodiment, a data storage system is constructed using one computer 10, and in the second embodiment, a data storage system is constructed using multiple computers constituting multiple clients 40 and a server 50.

[0110] In contrast, the functions realized by these computers may be realized by a dedicated computer provided by configuring a processor with one or more dedicated hardware logic circuits. Alternatively, the functions may be realized by one or more dedicated computers configured by combining a processor and memory programmed to execute one or more functions with a processor configured with one or more hardware logic circuits. Furthermore, a computer program for realizing each of the above functions may be stored in a computer-readable non-transitory tangible recording medium as instructions to be executed by a computer. Furthermore, the method for realizing each of the above functions does not necessarily need to include software, and all of the functions may be realized using one or more pieces of hardware.

[0111] In addition, multiple functions possessed by one component in the above embodiments may be realized by multiple components, or one function possessed by one component may be realized by multiple components. Also, multiple functions possessed by multiple components may be realized by one component, or one function realized by multiple components may be realized by one component. Also, part of the configuration of the above embodiments may be omitted. Also, at least part of the configuration of the above embodiments may be added to or substituted for the configuration of another of the above embodiments.

[0112] In addition, the technology disclosed herein can also be realized in various forms, such as the data storage system described above, a system that includes a data storage system as a component, a program for causing a computer to function as a data storage system, a non-transient physical recording medium such as a semiconductor memory on which this program is recorded, and a data storage method.

[0113] [Technical idea disclosed in this specification] [Item 1] an outlier detection unit configured to detect outliers that are isolated from a main data group and occur infrequently from the input data; A learning unit configured to input main data obtained by excluding the outliers from the input data into a data generation model using a machine learning model to learn the data generation model; a data storage unit configured to store the data generation model trained by the learning unit and the outlier detected by the outlier detection unit; A data storage system comprising:

[0114] [Item 2] a plurality of clients each including the outlier detection unit and the learning unit, each configured to transmit parameters of the data generation model obtained by the learning unit and the outliers detected by the outlier detection unit; 2. The data storage system according to item 1, further comprising a server configured to generate and update the data generation model using the parameters transmitted from the multiple clients, and to store the outliers transmitted from the multiple clients, as the data storage unit.

[0115] [Item 3] 3. The data storage system according to claim 1, wherein the data storage unit is configured to store the number of input data and the number of outliers.

[0116] [Item 4] an input unit capable of inputting a ratio between the number of the outliers detected by the outlier detection unit and the number of the main data; The data storage system according to any one of items 1 to 3, wherein the outlier detection unit is configured to detect the outliers such that the ratio between the number of the outliers and the number of the main data becomes the ratio input from the input unit.

[0117] [Item 5] 5. The data storage system of item 4, further comprising a display unit configured to display the accuracy of data reconstruction by the data generation model trained by the learning unit and / or the number of outliers detected by the outlier detection unit.

[0118] [Item 6] The data storage system according to any one of items 1 to 5, wherein the data generation model learned by the learning unit and the outliers detected by the outlier detection unit are stored for each preset period together with time information representing the period.

[0119] [Item 7] The outlier detection unit each time the learning unit performs the learning based on the input data, a reconstruction accuracy of data by the learned data generation model is calculated, and when the reconstruction accuracy is less than a predetermined threshold, the input data used in the learning is determined to be an outlier, and the learning unit is caused to perform the learning based on main data obtained by excluding the outlier from the input data. The data storage system according to any one of items 1 to 6, configured as described above. [Explanation of symbols]

[0120] 1...data storage system, 2...input section, 4...display section, 12...outlier detection section, 14...learning section, 16...data storage section, 40...client, 50...server.

Claims

1. an outlier detection unit (12) configured to detect outliers that are isolated from a main data group and occur infrequently from input data; A learning unit (14) configured to input main data obtained by removing the outliers from the input data into a data generation model using a machine learning model to learn the data generation model; a data storage unit (16) configured to store the data generation model trained by the learning unit and the outliers detected by the outlier detection unit; A data storage system comprising:

2. The method further comprises: providing a plurality of clients (40) each including the outlier detection unit and the learning unit, and configured to transmit parameters of the data generation model obtained by the learning unit and the outliers detected by the outlier detection unit; The data storage system of claim 1, further comprising a server (50) configured to generate and update the data generation model using the parameters transmitted from the plurality of clients and to store the outliers transmitted from the plurality of clients, as the data storage unit.

3. The data storage system according to claim 1 or 2, wherein the data storage unit is configured to store the number of the input data or the main data and the number of the outliers.

4. an input unit (2) capable of inputting a ratio between the number of the outliers detected by the outlier detection unit and the number of the main data; 3. The data storage system of claim 1, wherein the outlier detection unit is configured to detect the outliers so that the ratio between the number of the outliers and the number of the main data becomes the ratio input from the input unit.

5. The data storage system of claim 4, further comprising a display unit (4) configured to display the accuracy of data reconstruction by the data generation model learned by the learning unit and / or the number of outliers detected by the outlier detection unit.

6. The data storage system of claim 1 or claim 2, configured to store the data generation model learned by the learning unit and the outliers detected by the outlier detection unit for each predetermined period, together with time information representing the period.

7. The outlier detection unit each time the learning unit performs the learning based on the input data, a reconstruction accuracy of data by the learned data generation model is calculated, and when the reconstruction accuracy is less than a predetermined threshold, the input data used in the learning is determined to be an outlier, and the learning unit is caused to perform the learning based on main data obtained by excluding the outlier from the input data.

3. The data storage system according to claim 1 or 2, configured as follows.

Citation Information

Patent Citations

  • Data compression device, data reproduction device, data compression method, data reproduction method and data transfer method

    JP2018061091A