DATA PROCESSING METHOD, DATA PROCESSING APPARATUS AND DATA PROCESSING PROGRAM

By calculating the variance covariance matrix eigenvalue of high-dimensional data and performing data correction, the problem of difficulty in removing high-dimensional data noise in the prior art is solved, and effective data correction and feature extraction are realized.

JP7672160B2Active Publication Date: 2025-05-07KYOTO UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022531772
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-15
Filing Date
2021-06-11
Publication Date
2025-05-07
Estimated Expiration
2041-06-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively remove noise from high-dimensional data, resulting in the inability to obtain correction data close to real data.

Method used

By calculating the eigenvalue of the variance covariance matrix of the observed data, and using the noise scanning method or the intersection matrix method to calculate and modify the eigenvalue, perform data correction, make the eigenvalue match, and keep the eigenvector of the covariance matrix unchanged.

Benefits of technology

Effectively reduce data noise and obtain corrected data close to real data, so that high-dimensional data analysis can more accurately extract the real characteristics and structure of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672160000026
    Figure 0007672160000026
  • Figure 0007672160000027
    Figure 0007672160000027
  • Figure 0007672160000028
    Figure 0007672160000028
Patent Text Reader

Abstract

The present invention is a data processing method that reduces noise from data having a large number of feature amounts that are the basis for analysis, and obtains corrected data that is close to the true data. In the present invention, a data processing method is provided for reducing noise from data having a large number of feature amounts that are the basis for analysis and obtaining corrected data that is close to the true data, wherein the eigenvalue of a variance-covariance matrix of the data matches a prescribed eigenvalue, and the data processing method further includes a data correction step for correcting the data so that the eigenvector of the variance-covariance matrix does not vary before and after correction. Furthermore, the data is corrected so that the statistical amount of the feature amounts of the data does not vary before and after correction.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a data processing method, a data processing device, and a data processing program. [Background technology]

[0002] In the above technical fields, when analyzing data with a large number of features, such as high-dimensional data, problems such as the curse of dimensionality due to the presence of noise arise. Patent documents 1 and 2 disclose techniques for correcting noise components using statistical methods such as principal component analysis, with the aim of reducing noise from data with a large number of features that are the basis of the analysis. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Florian Wagner, Dalia Barkley and Itai Yanai, "Accurate denoising of single-cell RNA-Seq data using unbiased principal component analysis", BioRxiv(2019):655365. [Non-Patent Document 2] David van Dijk et al., "Recovering Gene Interactions from Single-Cell Data Using Data Diffusion", Cell 174, 716-729, July 26, 2018, Elsevier Inc. [Non-Patent Document 3] Makoto Aoshima, Kazuyoshi Yata, "Statistical Methodology for High-Dimensional Data," Journal of the Japan Statistical Society, Vol. 43, No. 1, September 2013, pp. 123-150 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the techniques described in Patent Documents 1 and 2 above were unable to obtain corrected data close to the true data by reducing noise from the data itself, which has a large number of feature quantities and is the basis for analysis.

[0005] An object of the present invention is to provide a technique for solving the above-mentioned problems. [Means for solving the problem]

[0006] In order to achieve the above object, the method according to the present invention comprises: A data processing method using a computer for removing noise from high-dimensional observation data, comprising the steps of: Using a processor included in the computer, a first calculation step of calculating eigenvalues ​​by solving a characteristic equation of a variance-covariance matrix of the observation data; a second calculation step of calculating a modified eigenvalue from the observed data using a noise sweeping method or a cross data matrix method; a data correction step of correcting the observed data so that the eigenvalues ​​calculated in the first calculation step are equal to the corrected eigenvalues ​​calculated in the second calculation step, and further so that the eigenvectors of the variance-covariance matrix do not change before and after the correction; It is a data processing method for executing the above.

[0007] In order to achieve the above object, the program according to the present invention comprises: A data processing program executed by a computer to remove noise from high-dimensional observation data, comprising: a first calculation step of calculating eigenvalues ​​by solving a characteristic equation of a variance-covariance matrix of the observation data; a second calculation step of calculating a modified eigenvalue from the observed data using a noise sweeping method or a cross data matrix method; a data correction step of correcting the observed data so that the eigenvalues ​​calculated in the first calculation step are equal to the corrected eigenvalues ​​calculated in the second calculation step, and further so that the eigenvectors of the variance-covariance matrix do not change before and after the correction; It is a data processing program that causes a computer to execute the above.

[0008] In order to achieve the above object, the device according to the present invention comprises: A data processing device for removing noise from high-dimensional observation data, comprising: a first calculation unit that calculates eigenvalues ​​by solving a characteristic equation of a variance-covariance matrix of the observation data; a second calculation unit that calculates a modified eigenvalue from the observed data by using a noise sweeping method or a cross data matrix method; a data correction unit that corrects the observed data so that the eigenvalues ​​calculated by the first calculation unit match the corrected eigenvalues ​​calculated by the second calculation unit and so that the eigenvectors of the variance-covariance matrix do not change before and after the correction; The data processing device includes: Effect of the Invention

[0009] According to the present invention, it is possible to obtain corrected data close to the true data by reducing noise from the data itself having a large number of feature quantities that are the basis for analysis. [Brief description of the drawings]

[0010] [Figure 1] 4 is a flowchart showing the procedure of a data processing method according to the first embodiment. [Diagram 2] FIG. 11 is a diagram for explaining the positioning of a data processing method as a data noise reduction method according to the second embodiment. [Figure 3A] FIG. 1 is a diagram for explaining an overview of a data analysis method according to a base technology. [Figure 3B] FIG. 1 is a diagram for explaining an overview of a data analysis method according to a base technology. [Figure 3C] FIG. 1 is a diagram for explaining an overview of a noise sweeping method according to a base technology. [Figure 4A] FIG. 11 is a diagram for explaining an overview of a data analysis method including a data processing method according to a second embodiment. [Figure 4B] FIG. 11 is a diagram illustrating requirements for a data processing method according to a second embodiment. [Figure 5A] 10 is a flowchart showing the procedure of a data processing method according to a second embodiment, in which no parameter adjustment is performed. [Figure 5B] 10 is a flowchart showing the procedure of a data processing method for adjusting parameters according to the second embodiment. [Figure 5C] FIG. 11 is a diagram illustrating a parameter adjustment method according to the second embodiment. [Figure 6] FIG. 11 is a diagram illustrating a first estimation method of parameters according to the second embodiment. [Figure 7] FIG. 11 is a diagram illustrating a second estimation method of parameters according to the second embodiment. [Figure 8] FIG. 11 is a block diagram showing the configuration of a data processing device according to a second embodiment. [Figure 9] FIG. 13 is a diagram illustrating requirements for a data processing method according to the third embodiment. [Figure 10A] 13 is a flowchart showing the procedure of a data processing method according to the third embodiment, in which parameter adjustment is not performed. [Figure 10B] 13 is a flowchart showing the procedure of a data processing method for adjusting parameters according to the third embodiment. [Figure 11A] FIG. 13 is a diagram illustrating a method for generating first data for verifying a data processing method according to the third embodiment. [Figure 11B] 13A to 13C are diagrams illustrating verification results of the data processing method according to the third embodiment. [Figure 12A] FIG. 13 is a diagram illustrating an overview of preprocessing in a data processing method according to a fourth embodiment. [Figure 12B] 13 is a flowchart showing the procedure of a data processing method including preprocessing according to the fourth embodiment. [Figure 13A] FIG. 13 is a diagram illustrating an overview of preprocessing in a data processing method according to a fifth embodiment. [Figure 13B] 13 is a flowchart showing the procedure of a data processing method including preprocessing according to the fifth embodiment. [Figure 14A] FIG. 11B shows the noise-reduced data generated by the method shown in FIG. 11A and verified by multidimensional scaling (MDS). [Figure 14B] FIG. 11B is a diagram showing the state in which noise is reduced from the data generated by the method shown in FIG. 11A and verified using t-distribution stochastic neighbor embedding (tSNE). [Figure 14C] FIG. 11B is a diagram showing the state in which noise is reduced from the data generated by the method shown in FIG. 11A and verified by topological data analysis (Mapper). [Figure 15] 11 is a diagram illustrating a method for generating second data for verifying the data processing method according to the present embodiment. FIG. [Figure 16A] This figure shows how noise is reduced from the data generated by the method shown in Figure 15 and verified using principal component analysis (PCA), multidimensional scaling (MDS), t-SNE, and dimensionality reduction (UMAP). [Figure 16B] This figure shows how noise is reduced from the data generated by the method shown in Figure 15 and verified using principal component analysis (PCA), multidimensional scaling (MDS), t-SNE, and dimensionality reduction (UMAP). [Figure 16C] This figure shows how noise is reduced from the data generated by the method shown in Figure 15 and verified using principal component analysis (PCA), multidimensional scaling (MDS), t-SNE, and dimensionality reduction (UMAP). [Figure 17A] FIG. 13 is a diagram showing the results of hierarchical clustering and a heat map when noise is reduced from observed gene expression data of a single cell using this embodiment. [Figure 17B] FIG. 13 is a diagram showing the verification of the results of hierarchical clustering when noise is reduced from observed gene expression data of a single cell according to the present embodiment. [Figure 18A]FIG. 13 is a diagram verifying the effect of noise reduction according to the present embodiment when comparing gene expression values ​​between two cells and comparing statistics of each gene among all cells based on time-course gene expression data of a germ cell differentiation induction system from human ES cells. [Figure 18B] FIG. 13 is a diagram showing the effect of noise reduction according to the present embodiment when clustering verification was performed using the tSNE method based on time-course gene expression data of a germ cell differentiation induction system from human ES cells. [Figure 18C] This figure shows the effect of noise reduction by this embodiment in a tSNE diagram in which the expression values ​​of four major genes (SOX2, SOX17, TFAP2C, NANOS3) are shown in heat color based on time-course gene expression data of a germ cell differentiation induction system from human ES cells. [Figure 19A] FIG. 13 is a diagram showing the effect of noise reduction achieved by preprocessing in the fifth embodiment on single-cell gene expression data (21,852 genes (dimensions)) of mouse gastrulation-stage embryos (E7.5, E8.5) collected using 10X Chromium. [Figure 19B] FIG. 13 is a diagram showing a comparison between a distribution obtained by fluorescence-activated cell sorting (FACS), with TFAP2C gene expression on the horizontal axis and BLIMP1 gene expression on the vertical axis, and a distribution obtained by noise reduction preprocessing according to the fifth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. However, the components described in the following embodiment are merely examples, and are not intended to limit the technical scope of the present invention.

[0012] [First embodiment] A data processing method S100 according to a first embodiment of the present invention will be described with reference to Fig. 1. The data processing method S100 is a method for correcting data so as to reduce noise in the data.

[0013] As shown in Fig. 1, the data processing method S100 includes a data correction step S101. In the data correction step S101, the data is corrected so that the eigenvalues ​​of the variance-covariance matrix of the data match a predetermined eigenvalue (S111). Furthermore, in the data correction step S101, the data is corrected so that the eigenvectors of the variance-covariance matrix do not change before and after the correction (S113). Note that the processing of steps S111 and S113 may be performed in this order, or these processing may be performed in parallel.

[0014] According to this embodiment, it is possible to obtain corrected data close to the true data by reducing noise from the data itself that has a large number of feature quantities that are the basis of the analysis. Note that in this embodiment and the following embodiments, the data is corrected so that the eigenvalues ​​of the variance-covariance matrix match the predetermined eigenvalues ​​and the eigenvectors of the variance-covariance matrix do not change before and after the correction. However, the present invention is not limited to this, and data correction that reduces each error to less than 10% is also permitted.

[0015] [Second embodiment] Next, a data processing method according to a second embodiment of the present invention will be described. Generally, high-dimensional data with a large number of features, for example, observed data for gene expression data, contains noise compared to the true data. When the dimension d of the data becomes high, the so-called "curse of dimensionality" occurs, and the characteristics and structure of the true data cannot be seen in data analysis. That is, even minute noises in high-dimensional observed data have a large effect on distance calculation, making it impossible to measure distance accurately. In addition, in high-dimensional observed data, inconsistencies occur in the data existence space, and for example, the contribution rate of the principal component is calculated to be low in principal component analysis (PCA).

[0016] In the data processing method of this embodiment, by performing processing that satisfies the following requirements, noise is reduced from each data itself and an approximation value of the true data is obtained, which makes it possible to extract the feature amount of the true data.

[0017] One is to correct the data so that the eigenvalues ​​of the variance-covariance matrix of the data match the predetermined eigenvalues, so as to make the eigenvalues ​​match the noise sweeping method, thereby performing the main data correction and noise removal. This requirement is hereinafter referred to as R1.

[0018] The second is to preserve the data structure by correcting the data so that the eigenvectors of the variance-covariance matrix of the data do not change before and after the correction, and by matching the results of the principal component analysis except for the scale factor. This requirement is hereafter referred to as R2.

[0019] By correcting each piece of data to meet these two requirements, the noise in each piece of data is reduced and the true characteristics and structure of the data can be seen in data analysis.

[0020] The contents represented by the symbols used in this embodiment are summarized below in formula (1).

number

[0021] <Positioning of the data processing method according to this embodiment> FIG. 2 is a diagram for explaining the positioning of the data processing method as a data noise reduction method according to this embodiment.

[0022] The characteristics and structure of the true data appear accurately in the data analysis method based on the true data 210. For example, when gene expression data is considered, in the true data held by cells, the expression data can be separated into fine cell types and the structure can be accurately detected.

[0023] However, high-dimensional observation data acquired by experiments or the like contains observation errors (noise) with respect to the true data due to the influence of observation equipment, observation environment, etc. In the data analysis method 220 based on high-dimensional observation data containing noise, since noise is contained in each dimension, separation is not possible or incorrect separation is performed, making it impossible to see the structure or to accurately analyze it.

[0024] Therefore, by using the data noise reduction method S200 of this embodiment to generate corrected data that is close to the true data from high-dimensional observation data, it becomes possible to analyze the characteristics and structure of the true data in the data analysis method 230 based on the corrected data.

[0025] That is, the data processing method of this embodiment does not improve the data analysis results by adjusting the data analysis method, but rather reduces noise from the high-dimensional observation data itself to bring it closer to the true data, thereby making it possible to analyze the inherent characteristics and structure of the true data. Therefore, the data analysis results of data containing noise can be improved for any data analysis method.

[0026] <Prerequisite technology> FIG. 3A is a diagram for explaining an overview of a data analysis method 210 according to the base technology.

[0027] Figure 3A shows the results of different data analysis methods when noise-free observation data, i.e., true data, are obtained using ideal observation equipment and in an ideal observation environment.

[0028] In Figure 3A, PCA is principal component analysis, MDS is multidimensional scaling, tSNE is t-distribution stochastic neighbor embedding, and Mapper is a topological data analysis method. Each analysis method can obtain the characteristics and structure of the true data.

[0029] FIG. 3B is a diagram for explaining an overview of a data analysis method 220 according to the base technology.

[0030] FIG. 3B shows the results of different data analysis when observation data containing noise is acquired using an observation instrument and in an observation environment.

[0031] In Figure 3B, we can see that the true data structure can be analyzed by PCA, but the contribution rate 321, which is a feature of PCA, is underestimated. On the other hand, MDS, tSNE, and Mapper cannot obtain the features and structure of the true data due to the influence of noise.

[0032] The principle by which the features and structure of the true data cannot be obtained due to noise in high-dimensional observational data, as shown in Figure 3B, is shown in equation (2). In other words, as the dimension (d) increases, noise increases the distance of the observational data x from the true data (0).

number

[0033] Furthermore, when high-dimensional observation data containing noise is subjected to principal component analysis, the phenomenon in which the characteristics and structure of the true data cannot be obtained is shown in equation (3). In other words, the eigenvalues ​​of the result of principal component analysis are buried in noise.

number

[0034] FIG. 3C is a diagram for explaining an outline of the noise sweeping method according to the base technology.

[0035] Graph 310 in FIG. 3C shows the area where the eigenvalues ​​of the principal component analysis are buried in noise according to the above-mentioned formula (3). i ,γ(n=cd γ )) is the area under the line connecting (0,1) and (1,0).

[0036] Equation (4) shows the mathematical correction method of the eigenvalues ​​of the principal component analysis, and equation (5) shows the phenomenon improved by the noise sweeping method compared to equation (3).

number

number

[0037] Graph 320 in FIG. 3C illustrates the modification of the eigenvalues ​​of the principal component analysis (from solid to dashed lines), and graph 330 illustrates the reduction in the area where the eigenvalues ​​of the principal component analysis are buried in noise due to equation (5) above. i ,γ(n=cd γ )) is the area under the line connecting (0,1) and (1 / 2,0).

[0038] As described above, in the noise sweeping method, the eigenvalues ​​of the principal component analysis are mathematically corrected, thereby reducing the area in which the eigenvalues ​​of the principal component analysis are buried in noise by half.

[0039] For details of the noise sweeping method, please refer to Non-Patent Document 3. The effect of the noise sweeping method shown in FIG. 3C is shown in FIG. 5 (page 132) of Non-Patent Document 3.

[0040] <Data Processing Method of the Present Embodiment> The data processing method as the data noise reduction method of this embodiment will be described in detail below.

[0041] (overview) FIG. 4A is a diagram for explaining an overview of a data analysis method 230 including a data processing method according to this embodiment.

[0042] The data noise reduction method S200 is applied to observation data having noise due to the observation equipment and the observation environment to correct each data itself. In any analysis method based on corrected data corrected to reduce noise from the observation data, the characteristics and structure of the true data can be obtained. In particular, since the contribution rate 421 of the principal components PC1 and PC2 in the PCA analysis is almost equal to the contribution rate of the true data without noise, it is confirmed that the observation data has been corrected to the true data by the data noise reduction method S200 of this embodiment.

[0043] (Requirements for data correction) FIG. 4B is a diagram for explaining requirements 400 of the data processing method according to this embodiment.

[0044] In FIG. 4B, noise is removed from the observed data X by a data noise reduction method (also referred to as R1R2 processing) that satisfies requirements R1 and R2. ~ R1-R2 is generated.

[0045] Requirements R1 and R2 are expressed by the following equation (6).

number

[0046] Requirement R1 indicates the requirement for removing noise by correcting data so that the eigenvalues ​​match the noise sweeping method shown in the premise technology. That is, the eigenvalue λ of the variance-covariance matrix of the corrected data is set to a predetermined eigenvalue λ ~ In addition, the specified eigenvalue λ of requirement R1 is ~ is not limited to the modified eigenvalues ​​obtained by the noise sweeping method.

[0047] Requirement R2 shows the requirement for preserving the data structure by correcting the data to match the results of the principal component analysis except for the scale factor. In other words, it is sufficient that the eigenvectors of the variance-covariance matrix of the corrected data match the eigenvectors of the variance-covariance matrix of the data before correction, and that the eigenvectors of the variance-covariance matrix do not change before and after correction.

[0048] (Theoretical background of data revision) The theoretical background for reducing the noise of each data to meet the requirements of FIG. 4B is shown in Equation (7).

number

[0049] In addition, in formula (7), "Theorem" is a theorem that modifies the eigenvalues ​​of the variance-covariance matrix of the data to any specified eigenvalues ​​by arbitrarily selecting the components of the diagonal matrix L, and is a process based on requirement R1. Also, "Remark" is a note indicating that the modification of "Theorem" is also based on requirement R2. "Corollary" is the corollary of this embodiment derived from the theorem. Also, L in the data modification formula in "Theorem" indicates that the specified eigenvalues ​​of "Theorem" and "Corollary" can be replaced from the noise sweeping method to the cross data matrix method.

[0050] (Data noise reduction processing) The process for reducing noise in each data item according to the theoretical background of Equation 7 is shown in Equation (8).

number

[0051] Then, to further optimize the noise reduction, we set a parameter (threshold l δ ) is shown in equation (9).

number

[0052] <Data processing procedure> The data processing procedure according to this embodiment will be described with reference to FIGS. 5A and 5B.

[0053] (Procedure without parameter adjustment) FIG. 5A is a flowchart showing the procedure of a data processing method without parameter adjustment according to this embodiment.

[0054] In step S511, a matrix of observed data including noise is input. In step S513, a characteristic equation of the variance-covariance matrix is ​​solved to calculate eigenvalues. In step S515, a corrected eigenvalue is calculated using a noise sweeping method. In step S517, each data is corrected so as to satisfy requirements R1 and R2 according to this embodiment, and corrected data in which noise has been removed from the observed data is calculated. Note that in step S517, the parameter "l" is not adjusted. In step S519, the corrected data in which noise has been removed is output.

[0055] (Procedure for adjusting parameters) FIG. 5B is a flowchart showing the procedure of a data processing method for adjusting parameters according to this embodiment.

[0056] First, in step S521, the matrix of observed data including noise and parameter δ are input. Next, in step S513, the characteristic equation of the variance-covariance matrix is ​​solved to calculate eigenvalues. In step S515, corrected eigenvalues ​​are calculated using a noise sweeping method. Furthermore, in step S527, each data is corrected to satisfy requirements R1 and R2 according to this embodiment, and corrected data in which noise has been removed from the observed data is calculated, and parameter "l" is adjusted to "threshold l δ ". Then, in step S529, the corrected data from which the noise has been removed is output.

[0057] <How to adjust parameters> 5C is a diagram illustrating a parameter adjustment method 500 according to this embodiment. In the parameter adjustment method 500, the true dimensions of the principal components are estimated.

[0058] In Fig. 5C, the solid line graph shows the eigenvalues ​​versus principal component numbers when the noise sweeping method is not applied, and the dashed line graph shows the eigenvalues ​​versus principal component numbers when the noise sweeping method is applied.

[0059] In the parameter adjustment method 500, the parameter (threshold l δ ) The eigenvalues ​​of the principal component numbers below are the main information, but the parameters (threshold l δ ) or more are judged to be noise that is not the main information. Then, the parameter (threshold l δ ) is set to 0. This makes it possible to correct data while reducing the effects of noise.

[0060] Below, the parameters (threshold l δ We show how to estimate

[0061] (First estimation method for parameters) FIG. 6 is a diagram illustrating a first parameter estimation method according to the present embodiment. In the first parameter estimation method, the threshold value l δis determined using a function approximation of the eigenvalues. That is, the threshold l δ The sum of the subsequent eigenvalues ​​is determined by the statistics and dimensions of the data, and the threshold l is set. δ Determine.

[0062] In each of the upper graphs in FIG. 6, linear portions 610, 620, and 630 appear in the graph of the log of the eigenvalue versus the principal component number. These linear portions 610, 620, and 630 are obtained by adjusting the parameters (threshold l δ ) is the noise portion.

[0063] Therefore, in each of the above figures, the numbers of the principal components that fall into the linear parts 610, 620, and 630 from number 1 are set to the threshold value l δ By doing so, the data is corrected by removing the effect of noise on the eigenvalues. In other words, the parameters (threshold l δ Here, the approximating function is not limited to a straight line obtained by taking the logarithm, and an appropriate function can be selected depending on the data.

[0064] Such parameters (threshold l δ The first estimation method for can be expressed by equation (10).

number

[0065] (Second estimation method for parameters) FIG. 7 is a diagram illustrating a second parameter estimation method according to the present embodiment. In the second parameter estimation method, the threshold value l δ is determined using the statistics of the data and the relationship between the eigenvalues.

[0066] In this embodiment, the following formula (11) can be assumed. Here, X is the true data, and Y is the noise. Also, m is the number of bases of the true data, that is, the dimension that determines the nature of the data. Then, the variance of the noise σ2 Assume that is known.

number

[0067] Graph 720 in FIG. 7 shows a method for finding the variance of noise, which has been known up until now. Graph 720 shows feature quantities on the horizontal axis and variance on the vertical axis. For example, in the example of graph 720, the variances are mostly concentrated around "1.0". Variances around "1.0" are considered to be variances of noise only, and variances including data are farther from "1.0". Then, as in graph 730 in FIG. 7, the horizontal axis shows variance, and the vertical axis shows frequency of appearance. From this graph 730, the area around variance "1.0", where the frequency of appearance is at its maximum, is considered to be the variance of noise.

[0068] The noise variance σ obtained as shown in graphs 720 and 730 in FIG. 2 Using the above, the dimension m of the intersection of the graph 710 (threshold l δ ) is obtained.

[0069] <Configuration of data processing device> FIG. 8 is a block diagram showing the configuration of a data processing device 800 according to this embodiment.

[0070] 8, a CPU (Central Processing Unit) 810 is a processor for arithmetic control, and executes programs to realize data processing. A ROM (Read Only Memory) 820 stores fixed data and programs, such as initial data and programs. A network interface 830 controls communication with other devices, and controls, for example, the reception of observation data and the transmission of correction data and analysis results.

[0071] The RAM (Random Access Memory) 840 is a random access memory used by the CPU 810 as a work area for temporary storage. An area for storing data necessary for realizing the present embodiment is secured in the RAM 840. The observed data 841 is data in which noise is added to true data. The corrected data 842 is data in which noise is reduced from the observed data 841 by the data noise reduction method of the present embodiment. The PCA analysis result 843 is data of the analysis result obtained by performing a PCA analysis on the corrected data 842. The MDS analysis result 844 is data of the analysis result obtained by performing an MDS analysis on the corrected data 842. The tSNE analysis result 845 is data of the analysis result obtained by performing a tSNE analysis on the corrected data 842. The Mapper analysis result 846 is data of the analysis result obtained by performing a Mapper analysis on the corrected data 842. The transmitted / received data 847 is data transmitted / received via the network interface 830. The input / output data 848 is data input / output to / from a peripheral device via the input / output interface 860.

[0072] The storage 850 stores a database, various parameters, or the following data or programs required to realize this embodiment. The data noise reduction method storage unit 851 stores the data noise reduction method of this embodiment. The data noise reduction method storage unit 851 includes requirements R1, R2, and R3, a data correction algorithm including a calculation formula, a parameter estimation algorithm including a calculation formula, and a pre-processing algorithm including a calculation formula. The storage 850 stores the following programs. The data noise reduction processing program 852 is a program that reduces noise in this embodiment. The data correction module 853 is a module that generates correction data according to the data correction algorithm. The parameter estimation module 854 is a module that estimates parameters according to the parameter estimation algorithm. The pre-processing module 855 is a module that performs pre-processing according to the pre-processing algorithm.

[0073] The following peripheral devices are connected to the input / output interface 860. The display unit 861 displays the internal state and calculation results of the data processing device 800. The operation unit 862, which is composed of a keyboard, touch panel, and pointing device, inputs instructions and data to the data processing device 800. The voice input / output unit 863, which includes a speaker and microphone, realizes voice communication with the user. The storage medium 864 is capable of inputting and outputting data to and from the data processing device 800, and is also used for inputting observation data and outputting correction data.

[0074] In the configuration of the data processing device 800 in Figure 8, the observation data 841 in the RAM 840 and the like corresponds to a data acquisition unit that acquires data, and the configuration that generates corrected data 842 by executing a program or module stored in the storage 850 corresponds to a data correction unit that corrects the data so that the eigenvalues ​​of the variance-covariance matrix match predetermined eigenvalues ​​and further so that the eigenvectors of the variance-covariance matrix of the data do not change before and after the correction.

[0075] It should be noted that the RAM 840 and storage 850 in FIG. 8 do not show programs and data related to the general-purpose functions of the data processing device 800 or other feasible functions.

[0076] According to this embodiment, it is possible to obtain corrected data close to the true data by reducing noise from the data itself that has a large number of feature quantities as the basis for analysis. That is, by correcting the data so that the eigenvalues ​​of the variance-covariance matrix of the data match the predetermined eigenvalues ​​(R1) and further the eigenvectors of the variance-covariance matrix do not change before and after the correction (R2), it is possible to obtain an approximation of the true data by reducing noise from the data itself. Here, the noise sweeping method or the cross data matrix method is used as a method for calculating the corrected eigenvalue used for the predetermined eigenvalue.

[0077] Furthermore, in this embodiment, a parameter (threshold l) for selecting the true principal component and removing noise is δ ) can be estimated, and by removing the noise-only components of the principal component analysis, an approximation of the true data can be obtained.

[0078] [Third embodiment] Next, a data processing method according to a third embodiment of the present invention will be described. The data processing method according to this embodiment differs from the second embodiment in that the method further requires the preservation of a statistical law (law of large numbers) by matching averages (this requirement will be referred to as R3 below). The other configurations and operations are the same as those of the second embodiment, so the same configurations and operations are denoted by the same reference numerals and detailed descriptions thereof will be omitted.

[0079] <Data Processing Method of the Present Embodiment> (Requirements for data correction) Fig. 9 is a diagram for explaining requirements 900 of the data processing method according to this embodiment. Note that in Fig. 9, duplicated explanations of requirements similar to those in Fig. 4B will be omitted.

[0080] In FIG. 9, requirement R3 is added to FIG. 4B.

[0081] The requirements R1 to R3 of this embodiment, including the requirement R3, are shown in formula (12).

number

[0082] (Theoretical background of data revision) The theoretical background of data correction according to this embodiment is shown in equation (13).

number

[0083] (Data noise reduction processing) The process for reducing noise in each data item (also called R1R3 process) based on the theoretical background of Equation 13 is shown in Equation (14).

number

[0084] Then, to further optimize the noise reduction, we set a parameter (threshold l δ ) is shown in equation (15).

number

[0085] <Data processing procedure> The data processing procedure according to this embodiment will be described with reference to FIGS. 10A and 10B.

[0086] (Procedure without parameter adjustment) Fig. 10A is a flowchart showing the procedure of a data processing method without parameter adjustment according to this embodiment. In Fig. 10A, the same steps as in Fig. 5A are given the same step numbers, and duplicated explanations will be omitted.

[0087] In step S1017, each data is corrected so as to satisfy requirements R1, R2, and R3 according to this embodiment, and corrected data in which noise has been removed from the observed data is calculated. Note that in step S1017, the parameter "l" is not adjusted.

[0088] (Procedure for adjusting parameters) 10B is a flowchart showing the procedure of the data processing method for parameter adjustment according to this embodiment. First, in step S521, the matrix of observed data including noise and the parameter δ are input. Next, in step S513, the characteristic equation of the variance-covariance matrix is ​​solved to calculate eigenvalues. In step S515, the corrected eigenvalues ​​are calculated using the noise sweeping method.

[0089] Furthermore, in step S1027, each data is corrected so as to satisfy the requirements R1, R2, and R3 according to this embodiment, and corrected data in which noise has been removed from the observed data is calculated. The parameter “l” is adjusted to “threshold l δ ". Then, in step S529, the corrected data from which the noise has been removed is output.

[0090] <Verification of data processing method> FIG. 11A is a diagram illustrating a method 1110 for generating first data for verifying the data processing method according to this embodiment.

[0091] In Fig. 11A, 1,000 samples with two figure-8 structures in a certain three-dimensional space are embedded in 20,000 dimensions to generate high-dimensional data, which is used as pseudo-real data. By adding uniformly distributed noise to this real data, pseudo-observation data is generated. This observation data is generated multiple times while changing the noise, which is used as the first data.

[0092] FIG. 11B is a diagram for explaining a verification result 1120 of the data processing method according to this embodiment.

[0093] As shown in the left diagram of Figure 11B, samples from each observation data of the first data generated by the generation method of Figure 11A are used as input data to generate corrected data B using the data noise reduction method of the second embodiment that satisfies requirements R1 and R2, and corrected data A using the data noise reduction method of this embodiment that satisfies requirements R1, R2, and R3.

[0094] Then, the average value of the error between the corrected data A and the true data for each sample is calculated. This is repeated for multiple observations to calculate the average value of the error for each sample multiple times. The same calculation is performed for the corrected data B.

[0095] The lower right diagram of Fig. 11B shows an average value 1121 of the average value of the errors calculated based on corrected data A, and an average value 1122 of the average value of the errors calculated based on corrected data B. Here, at the top of the average values ​​1121 and 1122, a confidence interval that includes the average value by 99% is shown in the form of a "".

[0096] The difference between the upper limit of the "I"-shape of the average value 1121 and the lower limit of the "I"-shape of the average value 1122 indicates the difference in the degree of approximation to the true data between the correction data A and the correction data B. That is, it can be seen that the data noise reduction method of this embodiment that satisfies the requirements R1, R2, and R3 can generate correction data closer to the true data than the data noise reduction method of the second embodiment that satisfies the requirements R1 and R2.

[0097] Conversely, this proves that the deviation of the correction data from the true data in the data noise reduction method of the second embodiment that satisfies the requirements R1 and R2 is uniform and has no variation in the first place.

[0098] According to this embodiment, by adding the requirement R3, it is possible to reduce noise from the data itself having a large number of feature quantities serving as the basis for analysis and obtain correction data closer to the true data.

[0099] [Fourth Embodiment] Next, a data processing method according to a fourth embodiment of the present invention will be described. The data processing method according to this embodiment is different from the second and third embodiments in that normalization, which is a preprocessing for removing noise from each data, and an inverse transformation of the normalization, which is a postprocessing, are performed. In this embodiment, a data processing method (hereinafter referred to as γ processing) in which preprocessing is performed assuming that the noise distribution is a gamma distribution is shown. The normalization process is particularly effective for data including noise based on random sampling such as gene expression data. Since the other configurations and operations are the same as those of the second or third embodiment, the same configurations and operations are denoted by the same reference numerals and detailed descriptions thereof are omitted.

[0100] (Outline of the data processing method of this embodiment) FIG. 12A is a diagram for explaining the outline of the preprocessing in the data processing method according to this embodiment. In such preprocessing, a normalization process is performed based on the feature quantities of the data. For example, the normalization process includes a process of dividing by the square root of the average of the feature quantities.

[0101] First, problems in the data processing method according to this embodiment will be described with reference to graph 1210 in Fig. 12A. In particular, problems in single cell gene expression analysis (single cell RNA-seq; scRNA-seq) will be discussed.

[0102] The main noise in scRNA-seq is generated from observation errors based on random sampling, and the noise distribution is known to be a gamma distribution (a distribution in which the mean and variance are correlated) (see Grun et al., Validation of Noise Models for Single-Cell Transcriptomics, 2014 June, Nature Method, 11(6):637-40). Some genes in a cell determine the characteristics of the cell, but there are also genes that have little effect (are not important). Many of these genes that are not involved in determining the characteristics have functions that are involved in the survival of the cell itself (house keeping genes), and they are expressed uniformly at high levels in almost all cell types. In other words, the expression amount of house keeping genes has a very small variance in the true data, but because the expression amount is high, it has a high variance value in the observed data. In other words, in the data noise reduction method of this embodiment, the importance of data is defined based on the variance of principal component analysis, so that even features that do not affect the characteristics of the cell, such as house keeping genes, have a high importance. Therefore, noise remains even after noise reduction by the third embodiment.

[0103] Graph 1210 in Figure 12A is a graph with the average on the horizontal axis and the variance on the vertical axis based on 200 pseudo-observational data in which the expression levels of 10% of the genes in a single 20,000-dimensional gene expression data set have been perturbed and noise based on different gamma distributions has been added.

[0104] A graph 1210 plots important features 1211 to which perturbation has been applied and unimportant features 1212 to which only noise has been applied. If the importance is divided by δ (parameter) so that the upper part is important and the lower part is unimportant, the important features will be eliminated and the unimportant features (only noise) will be emphasized.

[0105] In this embodiment, based on the characteristics of the gamma distribution (variance is proportional to the mean), a normalization process is performed based on the features of the preprocessed data, and the data is corrected so that it approximately matches a rotated graph of graph 1210 in Fig. 12A. As a result, a graph 1220 equivalent to a rotated graph of graph 1210 in Fig. 12A is obtained, making it possible to separate important features from unimportant features (only noise).

[0106] The normalization process as the preprocessing is shown in equation (16).

number

[0107] When equation (16) is applied to a data matrix X containing real data and noise, it is expressed as equation (17).

number

[0108] (Procedure including pre-processing) Fig. 12B is a flowchart showing the procedure of a data processing method including pre-processing according to this embodiment. In Fig. 12B, the same steps as those in Figs. 5A, 5B, 10A, and 10B are given the same step numbers, and duplicated explanations are omitted.

[0109] In step S1212, the input data is subjected to the normalization process of this embodiment, and the matrix X is converted into a matrix Z1. In steps S1213 and S1215, the same processes as in steps S513 and S515 in FIG. 5A are performed using the converted matrix Z1. In step S1216, the noise variance θ and the corresponding parameter η1 are estimated. Note that the noise variance θ is calculated based on the σ estimated in the first parameter estimation. 2 The parameter η1 corresponds to the threshold l δ is equivalent to.

[0110] In step S1217, the threshold l in step S527 of FIG. δ" is replaced with "η1" to correct each data, and corrected data with noise removed from matrix Z1 is calculated. Note that the reason there is no +(average value of Z1) corresponding to requirement R3 is because (average value of Z1) is set to 0. Then, in step S1218, the inverse transformation of the normalization process is performed.

[0111] Since S1216 is independent of the eigenvalue calculation, any of the following orders may be used: S1213 → S1215 → S1216, or S1213 → S1216 → S1215, or S1216→S1213→S1215.

[0112] According to this embodiment, by adding pre-processing to unify the variance of noise, it is possible to reduce noise based on random sampling from the data itself that has a large number of features that are the basis of the analysis, and obtain corrected data that is closer to the true data. This is particularly effective for gene expression data.

[0113] [Fifth embodiment] Next, a data processing method according to a fifth embodiment of the present invention will be described. The data processing method according to this embodiment differs from the second and third embodiments in that normalization is performed as a pre-processing for removing noise from each data, and the inverse transformation of the normalization is performed as a post-processing. Furthermore, compared to the fourth embodiment, this embodiment shows a data processing method in which pre-processing is performed by estimating only the variance of noise (hereinafter referred to as variance processing), rather than performing pre-processing by assuming that the noise distribution is a gamma distribution. Since the other configurations and operations are the same as those of the second to fourth embodiments, the same reference numerals are used for the same configurations and operations, and detailed explanations thereof will be omitted.

[0114] Graph 1310 in FIG. 13A is a graph showing sample data used for testing the data processing method according to this embodiment, with the horizontal axis representing the average and the vertical axis representing the variance. The sample data is single cell gene expression analysis (single cell RNA-seq; scRNA-seq) data, and is data obtained from 10X Genomics. Single cell gene expression datasets. (https: / / support.10xgenomics.com / single-cell-gene-expression / datasets.). This sample data is sample data of 36,601 genes from 5,604 cells, and is 25,890-dimensional data when silent genes (genes with an expression level of 0 in all cells) are excluded.

[0115] In this sample data, like 1210 in FIG. 12A, noise remains in graph 1310 in FIG. 13A even after noise reduction according to the third embodiment. For example, if the importance is divided by δ (parameter) so that the upper part is important and the lower part is not important, important features will be eliminated and unimportant features (only noise) will be emphasized. In this embodiment, the data is corrected so that it approximately matches a rotated graph of graph 1310 in FIG. 13A by performing preprocessing to perform normalization processing based on the variance of the data. As a result, graph 1330 equivalent to a rotated graph of graph 1310 in FIG. 13A is obtained, so that important features and unimportant features (only noise) can be separated. Note that graph 1330 in FIG. 13A has the average on the horizontal axis and the estimated noise variance β on the vertical axis. i * This is a graph of the above.

[0116] The normalization process as pre-processing in this embodiment is shown in equation (18).

number

[0117] When equation (18) is applied to a data matrix X containing real data and noise, it is expressed as equation (19).

number

[0118] (Procedure including pre-processing) Fig. 13B is a flowchart showing the procedure of a data processing method including pre-processing according to this embodiment. In Fig. 13B, the same steps as those in Fig. 5A, Fig. 5B, Fig. 10A, Fig. 10B, and Fig. 12B are given the same step numbers, and duplicated explanations will be omitted.

[0119] In step S1312, the input data is subjected to the normalization process of this embodiment, and the matrix X is converted into a matrix Z2. In steps S1313 and S1315, the same processes as in steps S513 and S515 in FIG. 5A are performed using the converted matrix Z2. In step S1316, the parameter η2 is estimated. When the parameter η2 is less than the threshold l δ is equivalent to.

[0120] In step S1317, the threshold l in step S527 of FIG. δ Each data is corrected by replacing " with "η2", and corrected data in which noise has been removed from matrix Z2 is calculated. Then, in step S1318, the inverse transformation of the normalization process is performed.

[0121] Since S1316 is independent of the eigenvalue calculation, any of the following orders may be used. S1313 → S1315 → S1316, or S1313 → S1316 → S1315, or S1316→S1313→S1315.

[0122] According to this embodiment, by adding preprocessing based on the variance of noise, it is possible to reduce noise based on random sampling from the data itself that has a large number of features that are the basis of the analysis, and obtain corrected data that is closer to the true data. This is particularly effective for gene expression data.

[0123] The preprocessing of this embodiment provides the following advantages over the fourth embodiment. In the fourth embodiment, normalization was performed based on the average value of the features, so features with the same average value are normalized in the same way, and no difference appears. In contrast, the preprocessing of this embodiment takes into account the variation in each feature, so that features with the same average value are normalized differently, and more effective normalization is performed for each feature. This improves the effect of noise reduction for each feature. Furthermore, in this embodiment, the step of estimating θ (before S1216) in the fourth embodiment is eliminated, so the processing speed can be increased. EXAMPLES

[0124] <Verification of data processing method for the first data> (Verification by principal component analysis: PCA) In this embodiment, noise was reduced from the first data generated by the method shown in FIG. 11A, and the data was verified by principal component analysis (PCA).

[0125] Regarding PCA, verification is performed based on whether the contribution rate is close to the true data by removing noise from the observed data using the data processing method of this embodiment.

[0126] The contribution rate is calculated by the following formula (20).

number

[0127] The contribution rate calculation results are shown in Table 1. [Table 1]

[0128] As shown in Table 1, in the first four components of the principal components, the contribution rate of the true data without noise approaches the following order: when the data noise reduction method is not applied → when processing is performed without parameter adjustment → when processing is performed with parameter adjustment → when processing is performed with appropriate parameter adjustment. And when processing is performed with appropriate parameter adjustment (e.g., δ=0.005), it becomes almost the same as the true data. In other words, the corrected data when processing with appropriate parameter adjustment is performed is almost the same as the true data, which indicates that noise reduction has been performed reliably.

[0129] (Verification by multidimensional scaling: MDS) FIG. 14A is a diagram showing a state in which noise is reduced from the first data generated by the method shown in FIG. 11A and verified by multidimensional scaling (MDS).

[0130] The upper part of FIG. 14A shows an analysis result 1411 of real data and an analysis result 1412 of pseudo observed data with uniformly distributed noise superimposed thereon.

[0131] The lower part of Fig. 14A shows the results of analysis of corrected data obtained by applying the data noise reduction method of this embodiment to observed data. The analysis result 1413 without parameter adjustment, the analysis result 1414 with appropriate parameter adjustment, and the analysis result 1415 with appropriate parameter adjustment are shown. From the analysis result 1415 with appropriate parameter adjustment, it is clear that the characteristics and structure of the true data can be analyzed, and it is shown that noise has been accurately removed.

[0132] (Testing with t-distribution-based probabilistic neighbor embedding method: tSNE) FIG. 14B is a diagram showing a state in which noise is reduced from the first data generated by the method shown in FIG. 11A and verified using t-distribution stochastic neighbor embedding (tSNE).

[0133] The upper part of FIG. 14B shows an analysis result 1421 of real data and an analysis result 1422 of pseudo observed data with uniformly distributed noise superimposed thereon.

[0134] The lower part of Fig. 14B shows the results of analysis of corrected data obtained by applying the data noise reduction method of this embodiment to observed data. The analysis result 1423 without parameter adjustment, the analysis result 1424 with appropriate parameter adjustment, and the analysis result 1425 with appropriate parameter adjustment are shown. From the analysis result 1425 with appropriate parameter adjustment, it is clear that the characteristics and structure of the true data can be analyzed, and it is shown that noise has been accurately removed.

[0135] (Topological Data Analysis: Verification by Mapper) FIG. 14C is a diagram showing a state in which noise has been reduced from the first data generated by the method shown in FIG. 11A and verified by a topological data analysis (Mapper).

[0136] The upper part of FIG. 14C shows an analysis result 1431 of real data and an analysis result 1432 of pseudo observed data with uniformly distributed noise added thereto.

[0137] The lower part of Fig. 14C shows the results of analysis of corrected data obtained by applying the data noise reduction method of this embodiment to observed data. The analysis result 1433 without parameter adjustment, the analysis result 1434 with appropriate parameter adjustment, and the analysis result 1435 with appropriate parameter adjustment are shown. From the analysis result 1435 with appropriate parameter adjustment, it is clear that the characteristics and structure of the true data can be analyzed, and it is shown that noise has been accurately removed.

[0138] <Verification of data processing method for second data> FIG. 15 is a diagram illustrating a method 1500 for generating second data for verifying the data processing method of this embodiment.

[0139] In Figure 15, 500 gene expression data were generated to simulate five cell groups by adding 100 different perturbations to five different expression regions of the gene expression data of 18,333 actual genes. These (500 x 18,333) pieces of data were used as pseudo true data. Gamma-distributed noise was added to this true data to generate (500 x 18,333) pieces of pseudo observed data containing noise, which were used as the second data.

[0140] Hereinafter, various processes including the data noise reduction method of this embodiment are performed on this second data, and the results of multiple data analyses based on the data will be shown.

[0141] (Comparison of dimensionality reduction method and data processing method according to this embodiment) FIG. 16A shows the noise-reduced second data generated by the method shown in FIG. 15 and verified by principal component analysis (PCA), multidimensional scaling (MDS), t-SNE, and two-dimensional projection using a dimensionality reduction method (UMAP).

[0142] In Fig. 16A, column 1601 is the analysis result of true data generated by the method of Fig. 15, column 1602 is the analysis result of observed data including noise, column 1603 is corrected data (PCA30) using a dimensionality reduction method that deletes the 31st and subsequent principal components of PCA, column 1604 is the analysis result of corrected data that has been subjected to normalization processing, satisfies requirements R1 to R3, and has undergone parameter estimation, and has been subjected to a data noise reduction method in this embodiment.

[0143] In addition, row 1610 is the analysis result using principal component analysis (PCA), row 1620 is the analysis result using multidimensional scaling (MDS), row 1630 is the analysis result using t-type stochastic neighbor embedding (tSNE), and row 1640 is the analysis result using a dimensionality reduction method (UMAP).

[0144] The corrected data that has been subjected to the data noise reduction method of this embodiment can reveal the characteristics and structure of the true data using any data analysis method. In other words, by accurately reducing the influence of noise from the observation data in which noise is superimposed on the true data divided into five groups, the data can be divided into five groups using any data analysis method. This shows that the data noise reduction method of this embodiment can reduce noise from the observation data and return it to the true data.

[0145] (Comparison of data processing methods with and without preprocessing) Fig. 16B is a diagram showing a state in which noise is reduced from the second data generated by the method shown in Fig. 15, and verified by principal component analysis (PCA), multidimensional scaling (MDS), t-type stochastic neighbor embedding (tSNE), and two-dimensional projection by a dimensionality reduction method (UMAP). In Fig. 16B, the same components as those in Fig. 16A are given the same reference numbers, and duplicated explanations are omitted.

[0146] In Fig. 16B, column 1605 is the analysis result of true data generated by the method of Fig. 15, and column 1606 is the analysis result of observed data including noise. Column 1607 is the analysis result of corrected data that has been subjected to a data noise reduction method in this embodiment, where the normalization process is not performed, requirements R1 to R3 are satisfied, and parameters are estimated. Column 1608 is the analysis result of corrected data that has been subjected to a data noise reduction method (gamma processing), where the normalization process is performed, requirements R1 to R3 are satisfied, and parameters are estimated.

[0147] The corrected data subjected to the data noise reduction method without normalization of this embodiment could separate two groups, but could not separate the other three groups. However, the corrected data subjected to the data noise reduction method with normalization can know the characteristics and structure of the true data by any data analysis method. That is, by accurately removing noise from the observation data in which noise was added to the true data divided into five groups, the data could be separated into five groups by any data analysis method. This shows that the data noise reduction method with normalization of this embodiment could remove noise from the observation data and return it to the true data.

[0148] (Comparison of dimensionality reduction method and data processing method with preprocessing by gamma processing and distribution processing) Fig. 16C is a diagram showing a state in which noise is reduced from the second data generated by the method shown in Fig. 15, and verified by principal component analysis (PCA), multidimensional scaling (MDS), t-type stochastic neighbor embedding (tSNE), and two-dimensional projection by a dimensionality reduction method (UMAP). Note that in Fig. 16C, the same components as those in Fig. 16A and Fig. 16B are given the same reference numbers, and duplicated explanations are omitted.

[0149] In FIG. 16C, column 1651 is the analysis result of true data generated by the method of FIG. 15, and column 1652 is the analysis result of observed data including noise. Column 1653 is corrected data (PCA30) by a dimensionality reduction method that deletes the 31st and subsequent principal components of PCA. Column 1654 is the analysis result of corrected data that has been subjected to a data noise reduction method (γ processing) in which normalization processing has been performed assuming that the noise distribution is a gamma distribution, requirements R1 to R3 have been satisfied, and parameters have been estimated in the fourth embodiment. Column 1655 is the analysis result of corrected data that has been subjected to a data noise reduction method (variance processing) in which normalization processing has been performed by estimating only the noise variance, requirements R1 to R3 have been satisfied, and parameters have been estimated in the fifth embodiment.

[0150] The corrected data that has been subjected to the data noise reduction method of this embodiment can reveal the characteristics and structure of the true data using any data analysis method. In other words, by accurately reducing the influence of noise from the observation data in which noise is superimposed on the true data divided into five groups, the data can be divided into five groups using any data analysis method. This shows that the data noise reduction method including any normalization process of this embodiment can reduce noise from the observation data and return it to the true data.

[0151] <Verification of data processing methods using single-cell gene expression data> (Effects of hierarchical clustering) FIG. 17A shows the results of hierarchical clustering and the corresponding heatmaps when noise is reduced by the present embodiment from single-cell gene expression data observed from mouse implantation-stage embryos (embryos aged 4.5 to 6.5 days after fertilization; E4.5 to 6.5).

[0152] The upper part of Fig. 17A shows dendrograms 1711 and 1721, which show the arrangement of cells on the horizontal axis and the distance on the vertical axis. The lower part of Fig. 17A shows heat maps 1712 and 1722 of representative genes whose order on the horizontal axis corresponds to that of the dendrograms 1711 and 1721.

[0153] In Fig. 17A, the dendrogram 1711 on the left side shows the result of hierarchical clustering performed on single-cell gene expression data without applying the data noise reduction method of this embodiment. As can be seen from the heat map 1712, noise is included in the observed data, so the distance between cells is large, and for example, in region 1713, adjacent cells cannot be classified correctly. As a result, cell groups that are originally grouped into two groups (E6.5 EPI L and E6.5 EPI H) are mixed into one cell group (E6.5 EPI).

[0154] In Fig. 17A, the dendrogram 1721 on the right shows the result of hierarchical clustering of single-cell gene expression data after applying the data noise reduction method of this embodiment. As can be seen from the heat map 1722, since what was thought to be included in the observed data disappears by gamma processing, the distance between each cell becomes small, and for example, in the region 1723, adjacent cells can be correctly classified. Therefore, the cell groups that are grouped into two groups (E6.5 EPI L and E6.5 EPI H) are clearly separated, and the gene expressions at E6.5 are also considerably aligned, appearing as different cell groups.

[0155] Figure 17B shows a dendrogram (left) that shows an enlarged portion of the hierarchical clustering results when noise is reduced from observed single-cell gene expression data using this embodiment, and a scatter plot comparing gene expression levels in the cells indicated by the arrows, which have been subjected to the data noise reduction method of this embodiment and those which have not.

[0156] In Fig. 17B, graphs 1714 and 1715 show analysis results based on observed data without the data noise reduction method of the present embodiment. In Fig. 17B, graphs 1724 and 1725 show analysis results based on corrected data with the data noise reduction method of the present embodiment.

[0157] Graphs 1714 and 1724 are scatter plots with the gene expression of cell E4.5PE_MS004_031 on the horizontal axis and the gene expression of cell E4.5PE_MS004_028 on the vertical axis, and graphs 1715 and 1725 are scatter plots with the gene expression of cell E4.5PE_MS004_031 on the horizontal axis and the gene expression of cell E4.5PE_MS004_047 on the vertical axis.

[0158] In graphs 1714 and 1715 of the analysis results based on observational data to which the data noise reduction method of this embodiment is not applied, the distributions are similar due to the influence of noise, and it is difficult to distinguish whether the relationship between the two cells is one group or different groups. In contrast, in graphs 1724 and 1725 of the analysis results based on observational data to which the data noise reduction method of this embodiment is applied, graph 1724 shows that the cells are from the same group, while graph 1725 shows that the cells are from different groups, which can be clearly distinguished.

[0159] In this way, by applying the data noise reduction method of this embodiment to generate corrected data from which noise has been removed, it becomes possible to easily and accurately classify cells based on their gene expression data.

[0160] <Application to cell induction experimental data> Next, we will verify whether the gamma processing of this embodiment correctly improves the results of data analysis in other studies using time-course gene expression data from a system in which germ cell differentiation from human stem cells is induced.

[0161] The data used was single-cell gene expression data from a human primordial germ cell induction experiment published in "D Chen et al., Cell Reports, 2019." The data used contained 22,044 cells and 18,252 genes, and included time changes (ES (day -2) → iMeLC (day 0) → day 1 → day 2 → day 3 → day 4).

[0162] In the analysis method of this example, the variance was calculated after log2 scaling for each of the two sets of data before and after the gamma processing of this embodiment. The variances were ranked in descending order, and the fluctuations in the ranking of genes considered to be biologically significant or insignificant in this induction experiment were examined.

[0163] The analysis results are shown in Table 2. [Table 2]

[0164] In Table 2, it can be seen that by applying the gamma processing of this embodiment to the experimental data, the ranking of major genes increases, while the ranking of minor genes decreases, thereby correctly removing the effects of noise contained in the experimental data.

[0165] Fig. 18A is a diagram showing the effect of noise reduction according to the present embodiment when gene expression was compared between two cells based on time-course gene expression data of a germ cell differentiation induction system from human ES cells. Fig. 18A shows several known genes that are considered to be important in germ lineage development.

[0166] The left diagram of Figure 18A is a scatter plot in which gene expression values ​​of two cells are plotted on the vertical and horizontal axes. Graph 1811 is based on uncorrected (original) gene expression data that has not been subjected to gamma processing according to this embodiment. On the other hand, graph 1821 is based on gene expression data that has been subjected to gamma processing according to this embodiment.

[0167] In scRNA-seq, genes with low expression levels are not detected due to dropout noise. Therefore, in the gene expression data before correction, the expression values ​​of many genes are "0" and are discrete. On the other hand, by applying gamma processing, the expression values ​​near "0" become greater than "0", and it can be seen that the dropout noise has been corrected.

[0168] The right diagram of Figure 18A is a log-log scale scatter plot of the average expression value of each gene for all cells in the data set on the horizontal axis and the variance of the expression value on the vertical axis. Graph 1812 is based on uncorrected gene expression data that has not been subjected to gamma processing according to the present embodiment. On the other hand, graph 1822 is based on gene expression data that has been subjected to gamma processing according to the present embodiment.

[0169] In the gene expression data before correction, the variance of the non-major genes is large due to the influence of noise, so the variance of the major genes is relatively insignificant. In contrast, after gamma processing, the variance of the non-major genes is reduced and the variance of the major genes is relatively large, so it is possible to determine that the major genes are significant.

[0170] FIG. 18B is a diagram showing the effect of noise reduction according to this embodiment when clustering verification using the tSNE method was performed based on time-course gene expression data of a germ cell differentiation induction system from human ES cells.

[0171] Clustering verification 1813 in Fig. 18B is based on gene expression data before correction that has not been subjected to gamma processing according to this embodiment. On the other hand, clustering verification 1823 is based on gene expression data that has been subjected to gamma processing according to this embodiment. In the gene expression data before correction, the boundaries separating the clusters are hardly generated, but after gamma processing, clear clusters can be observed.

[0172] Figure 18C shows the effect of noise reduction by this embodiment in a tSNE diagram in which the expression of four major genes (SOX2, SOX17, TFAP2C, NANOS3) is shown in heat color based on time-course gene expression data of a germ cell differentiation induction system from human ES cells.

[0173] In FIG. 18C, clustering verifications 1814 to 1817 are based on gene expression data before correction that has not been subjected to gamma processing according to this embodiment. On the other hand, clustering verifications 1824 to 1827 are based on gene expression data that has been subjected to gamma processing according to this embodiment. In the gene expression data before correction, there are regions in which SOX17, TFAP2C, and NANOS3 are highly expressed in the upper right corner, but clustering based on the expression values ​​of each is not possible (see dashed lines). However, in the gene expression data that has been subjected to gamma processing, three high expression clusters of SOX17 and TFAP2C are formed as shown by dashed lines in clustering verifications 1825 to 1827, and in clustering verification 1827 of NANOS3, 1871 of these clusters is a high expression cluster, and clustering based on the expression values ​​of major genes can be observed.

[0174] FIG. 19A is a diagram showing the effect of noise reduction, which is verified by preprocessing to estimate only the noise variance of the fifth embodiment, on single-cell gene expression data (21,852 genes (dimensions)) of mouse gastrulation-stage embryos (E7.5, E8.5) collected using 10X Chromium.

[0175] In Figure 19A, we verified the noise reduction method using clustering (tSNE). Mouse embryos at E7.5 and E8.5 are in the middle of a complex developmental process in which a wide variety of cell types appear, but the result without noise reduction (1910) does not reflect this reality. On the other hand, in the result with noise reduction (1920), various clusters appear, and cell types that were not visible in the result without noise reduction (1810) become visible.

[0176] Figure 19B shows a comparison between a distribution obtained by fluorescence-activated cell sorting (FACS), with TFAP2C gene expression on the horizontal axis and BLIMP1 gene expression on the vertical axis, and a distribution obtained by noise reduction in the fifth embodiment, which has been preprocessed to estimate only the noise variance.

[0177] Graph 1930 is the distribution of the observation results by FACS. Graph 1940 is raw data of single cell gene expression analysis (single cell RNA-seq; scRNA-seq) collected using 10X Chromium. Graph 1950 is the distribution of data resulting from noise reduction that has been pre-processed to estimate only the noise variance of the fifth embodiment. By comparative verification with the distribution by FACS, it was confirmed that true cell expression information could be restored by noise reduction.

[0178] [Other embodiments] In this embodiment and the present example, the present invention has been described in relation to multidimensional data, particularly gene expression data, but the present invention can be used to analyze data containing various types of noise, and has the same effect. For example, when applied to the analysis of measurement data (e.g., meteorological data, celestial data, material data, molecular data, image data, voice data, etc.) on scales such as length, weight, temperature, force, luminosity, saturation, and pitch, the same effect is achieved. In addition, when applied to the analysis of statistical data (e.g., gene expression data, epigenetic data, communication data, personal information, questionnaire data, etc.) obtained by (partially) collecting characteristics of living organisms or substances, the same effect is achieved.

[0179] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the technical scope of the present invention. In addition, systems or devices that combine the separate features included in each embodiment in any way are also included in the technical scope of the present invention.

[0180] The present invention may be applied to a system consisting of multiple devices, or to a single device. Furthermore, the present invention is also applicable to a case where an information processing program for implementing the functions of the embodiments is supplied to a system or device and executed by a built-in processor. In order to implement the functions of the present invention with a computer, a program installed on a computer, a medium storing the program, a server for downloading the program, and a processor for executing the program are also included in the technical scope of the present invention. In particular, at least a non-transitory computer readable medium storing a program for causing a computer to execute the processing steps included in the above-mentioned embodiments is included in the technical scope of the present invention.

[0181] This application claims priority based on Japanese Patent Application No. 2020-102852, filed on June 15, 2020, the disclosure of which is incorporated herein in its entirety.

Claims

1. A data processing method using a computer for removing noise from high-dimensional observation data, comprising the steps of: Using a processor included in the computer, a first calculation step of calculating eigenvalues ​​by solving a characteristic equation of a variance-covariance matrix of the observation data; a second calculation step of calculating a modified eigenvalue from the observed data using a noise sweeping method or a cross data matrix method; a data correction step of correcting the observed data so that the eigenvalues ​​calculated in the first calculation step are equal to the corrected eigenvalues ​​calculated in the second calculation step, and further so that the eigenvectors of the variance-covariance matrix do not change before and after the correction; A data processing method that performs the following:

2. 2. The data processing method according to claim 1, wherein in said data correction step, said observation data is corrected by the following formula: ##EQU00021##

3. 2. The data processing method according to claim 1, wherein said data correction step further comprises correcting said observed data so that statistics of feature quantities of said observed data do not change before and after the correction.

4. 4. The data processing method according to claim 3, wherein in said data correction step, said observation data is corrected by the following formula: [0022]

5. In the data correction step, the threshold value l of the eigenvector of the variance-covariance matrix of the observed data is δ 5. The data processing method according to claim 1, further comprising the step of correcting the observed data so that a main portion determined by the eigenvalues ​​of the eigenvectors of the other portions coincide with each other, and the eigenvalues ​​corresponding to the eigenvectors of the other portions become zero.

6. The threshold value l δ The data processing method according to claim 5 , wherein is expressed by the following formula: [0023]

7. The threshold value l δ The data processing method according to claim 5 or 6, wherein is determined using a relationship between a statistical quantity of the observed data and an eigenvalue.

8. The threshold value l δ 7. A method according to claim 5 or 6, wherein is determined using a functional approximation of the eigenvalues.

9. The threshold value l δ The threshold value l is set so that the sum of the subsequent eigenvalues ​​is a value determined by the statistics and dimensions of the observed data. δ 9. The data processing method of claim 8, further comprising the step of:

10. Prior to the data correction step, 10. The data processing method according to claim 1, further comprising a pre-processing step of performing a normalization process on the data.

11. The data processing method according to claim 10 , wherein the normalization process is a process of dividing the feature amount of the observed data by the ½ power of the average.

12. 11. The data processing method according to claim 10, wherein the normalization process is a process of dividing the observed data by the 1 / 2 power of the variance of noise in the observed data.

13. A data processing program executed by a computer to remove noise from high-dimensional observation data, comprising: a first calculation step of calculating eigenvalues ​​by solving a characteristic equation of a variance-covariance matrix of the observation data; a second calculation step of calculating a modified eigenvalue from the observed data using a noise sweeping method or a cross data matrix method; a data correction step of correcting the observed data so that the eigenvalues ​​calculated in the first calculation step are equal to the corrected eigenvalues ​​calculated in the second calculation step, and further so that the eigenvectors of the variance-covariance matrix do not change before and after the correction; A data processing program that causes a computer to execute the following:

14. 14. The data processing program according to claim 13, wherein the data correction step further causes the computer to correct the observed data so that statistics of feature quantities of the observed data do not change before and after the correction.

15. A data processing device for removing noise from high-dimensional observation data, comprising: a first calculation unit that calculates an eigenvalue by solving a characteristic equation of a variance-covariance matrix of the observation data; a second calculation unit that calculates a modified eigenvalue from the observed data by using a noise sweeping method or a cross data matrix method; a data correction unit that corrects the observed data so that the eigenvalues ​​calculated by the first calculation unit match the corrected eigenvalues ​​calculated by the second calculation unit and so that the eigenvectors of the variance-covariance matrix do not change before and after the correction; A data processing device comprising:

16. The data processing apparatus according to claim 15 , wherein the data correction section further corrects the observed data such that statistics of the feature quantities of the observed data do not change before and after the correction.

Citation Information

Patent Citations

  • Picture quality improving device

    JP2002222416A

  • Noise eliminating method and noise eliminating device

    JP2014153098A

  • Image processor, method for processing image, and program

    JP2019101686A