Data processing apparatus and data processing method

By updating the Gram matrix using SVD and calculating the covariance matrix in SVD form, the data processing apparatus efficiently computes the Mahalanobis distance, addressing computational intensity and numerical stability issues in high-dimensional feature spaces.

JP7693133B2Active Publication Date: 2025-06-16MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024560660
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-02-10
Publication Date
2025-06-16
Estimated Expiration
2043-02-10

AI Technical Summary

Technical Problem

The calculation of the Mahalanobis distance in data processing apparatuses requires the inverse matrix of the covariance matrix, which is computationally intensive and prone to numerical errors, especially when dealing with high-dimensional feature spaces.

Method used

The data processing apparatus updates the Gram matrix using singular value decomposition (SVD) in the learning phase and calculates the covariance matrix in SVD form, allowing for efficient calculation of the Mahalanobis distance while addressing numerical stability issues.

Benefits of technology

This approach reduces the computational burden and minimizes numerical errors, enabling efficient and accurate calculation of the Mahalanobis distance in high-dimensional feature spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693133000069
    Figure 0007693133000069
  • Figure 0007693133000070
    Figure 0007693133000070
  • Figure 0007693133000071
    Figure 0007693133000071
Patent Text Reader

Abstract

A data processing device according to the present disclosure comprises a processing circuit (120). The processing circuit (120) sequentially updates a Gram matrix in the form of an SVD in a training phase, and in Finalization of the training phase, the processing circuit (120) calculates a variance-covariance matrix in the form of an SVD on the basis of the SVD of the Gram matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed technology relates to a data processing apparatus and a data processing method.

Background Art

[0002] Data processing apparatuses are utilized, for example, in the field of machine learning. The problems handled by machine learning can be broadly classified into supervised learning and unsupervised learning. One example of supervised learning is the problem of predicting categories, i.e., "classification". Also, in unsupervised learning, there is the problem of finding groups, i.e., "clustering". Data processing apparatuses are used as artificial intelligence apparatuses for "classification" or as artificial intelligence apparatuses for "clustering".

[0003] As artificial intelligence for performing "classification" or "clustering" on images, for example, artificial neural networks such as CNN (Convolutional Neural Network) have achieved great results. An artificial neural network generates image feature amounts from image data. Image feature amounts are vector quantities that can be represented as vectors in a feature space. The data processing apparatus determines similarity or degree of deviation based on a distance definable in this feature space, for example, the Mahalanobis distance. Both similarity and degree of deviation are important quantities required in "classification" or "clustering".

[0004] The techniques of classification and clustering in machine learning are applied to anomaly detection for detecting anomalies from images. More specifically, anomaly detection is performed based on the "Mahalanobis distance" considering the estimation result of a normal distribution, assuming that the occurrence probability of samples belonging to a certain class in the feature space can be represented by a normal distribution. For example, Non-Patent Document 1 discloses a technique of applying a trained general-purpose CNN to anomaly detection using a method called Patch Distribution Modeling (sometimes simply referred to as "PaDiM").

Prior Art Documents

Non-Patent Literature

[0005]

Non-Patent Literature 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] The Mahalanobis distance is a distance defined by the covariance matrix (Σ). If the dimension of the feature space is N f then the covariance matrix (Σ) is a matrix of size N f ×N f The covariance matrix (Σ) is updated based on the training data in the learning phase. In the inference phase, the calculation of the Mahalanobis distance usually requires the inverse matrix of the covariance matrix (Σ), which is a matrix of size N f ×N f As algorithms for updating the covariance matrix (Σ) in the learning phase, singular value decomposition and an affinity-based algorithm, more specifically, a low-rank approximation based on singular value decomposition and an affinity-based algorithm, are required.

[0007]

Means for Solving the Problems

[0008] The data processing apparatus according to the disclosed technology includes a processing circuit. The processing circuit sequentially updates the Gram matrix in the form of SVD in the learning phase, and the processing circuit calculates the covariance matrix in the form of SVD based on the SVD of the Gram matrix described later in the Finalization of the learning phase.

Advantages of the Invention

[0009] The data processing apparatus according to the disclosed technology has the above configuration, and the algorithm for updating the scatter matrix (Σ) in the learning phase has an affinity with singular value decomposition and also has an affinity with low-rank approximation based on singular value decomposition. Accordingly, the data processing apparatus and the data processing method according to the disclosed technology can receive the benefits of singular value decomposition and the benefits of low-rank approximation based on singular value decomposition in the learning phase.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Embodiments for Carrying Out the Invention

[0011] 《Introduction Part 1》 The image feature amount (x) dealt with by the present disclosure technology is given as a vertical vector (column vector) as follows. TIFF0007693133000001.tif14166 Here, N f represents the length of the feature amount. The subscript "f" in N f is derived from the first letter of "feature". N d represents the number of learning data. In this specification, the learning data is hereinafter simply referred to as "data". N d The subscript "d" in is derived from the first letter of "data". Also, the variable c is derived from the first letter of "column" meaning a column (see Equation (2) described later). For a plurality of images belonging to a certain class, a data matrix (X) formed by arranging the respective image feature amounts is given by the following equation. TIFF0007693133000002.tif22166 In this specification, unless otherwise specified, the number of data is sufficient, and N d > N f is assumed.

[0012] The expected value (μ) of the image feature amounts belonging to the class is represented by the following equation. TIFF0007693133000003.tif36166 Here, the function E() represents the expected value. Generally, the expected value and the average value are different concepts, but in this case, the expected value (μ) of the image feature amounts is equal to the average (hereinafter referred to as the "average feature amount") of the N d image feature amounts belonging to the class.

[0013] The variance-covariance matrix (Σ) for the data matrix (X) is given by the following equation using μ. TIFF0007693133000004.tif38166 Here, the superscript "T" that appears in Equation (4) represents transposition. Note that the covariance matrix is sometimes simply referred to as the "covariance matrix". Also, XX that appears in the first term on the right side of Equation (4) T is called the Gram matrix. Also, the expected value of the Gram matrix, that is, E(XX in the first term on the right side of Equation (4) T ) is called the correlation matrix. The covariance matrix (Σ) can be derived by transforming Equation (4), but it can also be expressed using the deviation vector from the mean (x c -μ). TIFF0007693133000005.tif40166 Here, if we redefine the deviation vector (x c -μ) in the c-th column as y c , Equation (5) can be transformed as follows. TIFF0007693133000006.tif40166 The second row of Equation (6) represents that the covariance matrix (Σ) can be expressed as QQ T , that is, the covariance matrix (Σ) is a positive semi-definite matrix.

[0014] To measure the similarity or degree of deviation of the target sample (x target ) using the measurement results of the normal distribution, the Mahalanobis distance (d M ) is often used. The Mahalanobis distance (d M ) is given by the following equation using the covariance matrix (Σ). TIFF0007693133000007.tif27166 As shown in Equation (7), the Mahalanobis distance (d M ) is a distance that can be defined when the inverse matrix (Σ -1 ) of the covariance matrix (Σ) exists, that is, when the covariance matrix (Σ) is positive definite. The Mahalanobis distance (d M ) shown from Equation (1) to Equation (7) is a distance defined in the feature space of dimension N f .

[0015] The disclosed technology is interested in how the dispersion-covariance matrix (Σ) is updated when the data belonging to a certain class increases from N d to N d + 1, and clarifies this. When the dispersion-covariance matrix (Σ) when the data increases to N d + 1 is given in the singular value decomposition form, the inverse matrix (Σ M ) of the dispersion-covariance matrix (Σ) for calculating the Mahalanobis distance (d -1 ) can be easily obtained.

[0016] 《Introduction Part 2》 It is known that any matrix can be represented by its singular values and singular vectors. The form in which a matrix is decomposed into singular values and singular vectors is called singular value decomposition. The singular value decomposition of a certain p×q matrix (Z≠0) is represented as follows. TIFF0007693133000008.tif26166 Here, r represents the rank of Z. {σ1,..., σ r} appearing in Equation (8) are the singular values of Z respectively. S is a diagonal matrix with the singular values {σ1,..., σ r} as components. {u1,..., u r} appearing in Equation (8) are called the left singular vectors corresponding to the singular values {σ1,..., σ r}. Also, {v1,..., v r} appearing in Equation (8) are called the right singular vectors corresponding to the singular values {σ1,..., σ r}. The left singular vectors and the right singular vectors are collectively called singular vectors.

[0017] When the singular value decomposition of a certain p×q matrix (Z≠0) is given, research has been done on obtaining the singular value decomposition of Z + AB T . For example, the following non-patent literature discloses the algorithm. Matthew Brand, "Fast Low-Rank Modifications of the Thin Singular Value Decomposition", MERL Technical Report, TR2006-059, May 2006. Z + AB T The algorithm for obtaining the singular value decomposition of Z + AB can be described as a program and used as a function in a function library. In this specification, the function for obtaining the singular value decomposition of Z + AB T shall be represented as "IncrSVD" as follows. TIFF0007693133000009.tif30166 Here, "Incr" is derived from the first four characters of "Incremental" which means sequential, and "SVD" is derived from the initial letters of "Singular Value Decomposition" which means singular value decomposition.

[0018] 《Introduction Part 3》 When a certain matrix (Z, Z ≠ 0) can be represented by singular value decomposition (see Equation (8)), its Moore-Penrose type generalized inverse matrix (also called the pseudo-inverse matrix, hereinafter simply referred to as the "generalized inverse matrix") is given by the following equation. TIFF0007693133000010.tif46166 When Z is a regular matrix (p = q), the generalized inverse matrix of Z (Z - ) is the inverse matrix of Z (Z -1 ). While the inverse matrix is only defined for regular matrices, the generalized inverse matrix is defined for non-zero matrices. However, to calculate the generalized inverse matrix, the rank must be known. Note that since a vector is also a type of matrix, the generalized inverse matrix is defined for vectors as well.

[0019] In general, calculations handled in physics and engineering use observational data obtained from measuring devices and sensors, so calculation errors are always included. This is no exception in the technical field to which the data processing apparatus 100 according to the disclosed technology belongs. Therefore, when calculating the singular value decomposition for a matrix calculated based on observational data, all singular values (σ i、 i ranges from 1 to N f to natural numbers) become positive in numerical calculations. When a singular value that is originally 0 becomes non-zero due to numerical calculation errors, calculating the inverse matrix (including the generalized inverse matrix) of that matrix as it is will result in unrealistic values due to 1 / σ i .

[0020] To determine the rank of the p×q matrix Z obtained from measurement data, a method is adopted in which first a provisional rank (l = min(p,q)) is set and the following singular value decomposition is calculated. TIFF0007693133000011.tif28166 Next, for the value of the last singular value in the rank determination method, it is examined from which singular value it can be approximated to 0. For the specification of the allowable error of the rank of the matrix, for example, the smallest limit value that can be handled by a computer (referred to as "machine epsilon") is used. TIFF0007693133000012.tif9166 When the rank of Z is r, (Z) r is given by the following formula. TIFF0007693133000013.tif10166 The same method as the creation of (Z) defined in formula (13) r includes "low-rank approximation based on singular value decomposition". The generalized inverse matrix for the low-rank approximated matrix is referred to as the "rank-constrained generalized inverse matrix". The rank-constrained generalized inverse matrix is also referred to as the rank-constrained pseudo-inverse matrix.

[0021] Z and (Z) r have the following relational expression. TIFF0007693133000014.tif39166 Z and (Z) r The error between them is given by the following equation. TIFF0007693133000015.tif12166 Here, the symbol appearing on the left side of Equation (15) is the Frobenius norm or the Euclidean norm.

[0022] The rank-constrained generalized inverse matrix is a generalized inverse matrix using a low-rank approximation based on singular value decomposition. The low-rank approximation based on singular value decomposition has the same properties (drawbacks and advantages) as general approximations. The drawback of the approximation is that an error occurs with the true value. The advantage of the approximation is that the information can be slimmed down to the necessary amount. The low-rank approximation based on singular value decomposition also has the drawback that an error occurs with the true value and the advantage that the information can be slimmed down to the necessary amount.

[0023] Many of the matters described in Introduction 3 are cited from Reference 1 shown below. Also, the proof leading to the error shown in Equation (15) is shown in Reference 1. Reference 1: Kenichi Kanaya, "Linear Algebra Seminar, Projection, Singular Value Decomposition, Generalized Inverse Matrix", Kyoritsu Shuppan, ISBN978-4-320-11340-4.

[0024] Embodiment 1. The data processing device 100 according to Embodiment 1 shows the minimum necessary configuration required for the data processing device 100 according to the present disclosed technology. FIG. 1 is a hardware configuration diagram showing the hardware configuration of the data processing device 100 according to Embodiment 1.

[0025] FIG. 1A is a hardware configuration diagram showing the hardware configuration when the functions of the data processing apparatus 100 according to Embodiment 1 are executed by dedicated hardware. As shown in FIG. 1A, the hardware configuration when executed by dedicated hardware includes an input interface 110, a processing circuit 120, and an output interface 130. The processing circuit 120 may be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC, an FPGA, or a combination thereof. Each function of the data processing apparatus 100 may be realized by a separate processing circuit 120 or may be realized collectively by one processing circuit 120.

[0026] FIG. 1B is a hardware configuration diagram showing the hardware configuration when the functions of the data processing apparatus 100 according to Embodiment 1 are executed by software. As shown in FIG. 1B, the hardware configuration when executed by software includes an input interface 110, a processor 122, a memory 124, and an output interface 130. The processor 122 is generally also referred to as a CPU (Central Processing Unit), a central processing unit, a processing unit, an arithmetic unit, a microprocessor, a microcomputer, or a DSP (Digital Signal Processor).

[0027] When the hardware configuration of the data processing apparatus 100 is as shown in FIG. 1B, each function of the data processing apparatus 100 is realized by software, firmware, or a combination of software and firmware. The software and firmware are described as programs and stored in the memory 124. The processor 122 reads and executes the programs stored in the memory 124 to realize each function. That is, the data processing apparatus 100 includes a memory 124 for storing a program that, when executed by the processor 122, causes each processing step to be executed as a result. Also, these programs can be said to cause the processor 122 to execute the procedures and methods of the data processing apparatus 100 (the data processing method according to the present disclosure technology). Here, the memory 124 may be a non-volatile or volatile semiconductor memory such as a RAM, ROM, flash memory, or EPROM. Also, the memory 124 may include a disk such as a magnetic disk, flexible disk, optical disk, compact disk, mini disk, or DVD. Further, the memory 124 may be in the form of an HDD (Hard Disk Drive) or SSD (Solid State Drive).

[0028] Note that each function of the data processing apparatus 100 according to Embodiment 1 may be partially realized by dedicated hardware and the rest by software or firmware. As described above, the data processing apparatus 100 according to Embodiment 1 realizes each of the above functions by hardware, software, firmware, or a combination thereof.

[0029] The technical feature of the data processing apparatus 100 according to the present disclosure technology is that the scatter-covariance matrix (Σ) is not stored or used as a single matrix having a size of N f ×N f That is, it is not stored or used in that form. Since the scatter-covariance matrix (Σ) shown in Equation (6) is a positive semi-definite matrix, the scatter-covariance matrix (Σ) can be transformed as follows. TIFF0007693133000016.tif12166 Next, assume that the matrix Q appearing in Equation (16) is represented in the singular value decomposition form. TIFF0007693133000017.tif29166 In this specification, the singular values are arranged in descending order, and it is assumed that σ1 ≧ σ2 ≧ … ≧ σ NF is satisfied. Generally, sigma (especially "σ 2 ") is often used as a symbol representing variance. However, in this specification, as described above, sigma represents the singular value. From the properties of the singular vectors, the following equations hold for U and V appearing in Equation (17). TIFF0007693133000018.tif27166 Here, I appearing in Equation (18) is the identity matrix of size N f ×N f .

[0030] By substituting the singular value decomposition shown in Equation (17) into Equation (16), the covariance matrix (Σ) can be decomposed as follows. TIFF0007693133000019.tif31166 As shown in Equation (19), {σ1 2 , σ2 2 , …, σ Nf 2} are the singular values of the covariance matrix (Σ), respectively.

[0031] From the property shown in Equation (18) and Equation (19), the inverse matrix (Σ -1 ) of the covariance matrix (Σ) is given by the following equation. TIFF0007693133000020.tif37166 As described above, in the case of a regular matrix, the generalized inverse matrix coincides with the inverse matrix. In the derivation of Equation (20), it is assumed that all the singular values of the covariance matrix (Σ) are non-zero. The method for dealing with the case where the covariance matrix (Σ) is not full rank will become clear from the following explanation. In this specification, the form in which a matrix is decomposed and represented as a product of matrices is referred to as the "decomposition form". Singular value decomposition is a type of decomposition form and is a special form. What is given on the right side of Equation (20) is also a singular value decomposition. Thus, if the covariance matrix (Σ) is given by singular value decomposition, the problem that it appears to be full rank due to numerical errors (see Introduction Part 3) can be addressed. In this specification, singular value decomposition may also be referred to as SVD.

[0032] The data processing apparatus 100 according to the disclosed technology uses the decomposition form given on the right side of Equation (20) to calculate the Mahalanobis distance (d M ). TIFF0007693133000021.tif17166 As shown in Equation (21), in order to measure the similarity or degree of deviation of a certain sample (x target ), instead of the Mahalanobis distance (d M ), the square of the Mahalanobis distance (d M ), i.e., d M 2 ) can be used without any problem. Also, as shown in Equation (21), the feature amount (x target ) of the target can also be represented as a deviation vector (y target = x target - μ) defined as the deviation from the average feature amount (μ).

[0033] (Regarding the update method when the data increases) Figure 2 is a block diagram that schematizes the function used by the data processing apparatus 100 according to the disclosed technology. More specifically, Figure 2 represents the function IncrSVD() shown in Equation (9) as a block diagram.

[0034] FIG. 3 is a block diagram schematically showing a data update method used by the data processing apparatus 100 according to Embodiment 1. More specifically, FIG. 3 shows the case where the number of data increases from N d to N d +1, and is a block diagram schematically showing a data update method performed by the data processing apparatus 100 according to Embodiment 1. When the data increases, it is ideal but not essential that the matrices U, S, and V, which are in the singular value decomposition form of the scatter matrix (Σ), can be directly updated. As shown in FIG. 3, when the data increases, the data processing apparatus 100 according to Embodiment 1 updates the SVD of the Gram matrix using the first IncrSVD(), and obtains the SVD of the scatter matrix (Σ) from the SVD of the correlation matrix using the second IncrSVD(). In FIG. 3, the first IncrSVD() is shown as a vertically long block, and the second IncrSVD() is shown as a horizontally long block.

[0035] When the number of data is N d , assume that the SVD of the Gram matrix is given as follows. TIFF0007693133000022.tif11166 Here, the subscript “0” attached to the matrix appearing on the right side of Equation (22) is merely a number for identifying which matrix the SVD is for. When the number of data increases by N d +1, the Gram matrix is updated using the first IncrSVD(). TIFF0007693133000023.tif22166 Here, the subscript “1” attached to the matrix appearing on the left side of Equation (23) is, like “0” in Equation (22), a number for identifying which matrix the SVD is for.

[0036] The relationship between the Gram matrix and the correlation matrix is described in Equation (4). That is, the expected value of the Gram matrix is the correlation matrix. Based on this relationship, the SVD of the correlation matrix is given as follows. TIFF0007693133000024.tif24166

[0037] When x is newly added to the data nd+1 the average feature amount (μ|N d +1) is updated as follows. TIFF0007693133000025.tif15166 The process shown in Equation (25) is referred to as "update of average feature amount".

[0038] Finally, the SVD of the covariance matrix (Σ) when the number of data is N d +1 is given as follows using the updated average feature amount (μ|N d +1) and the second IncrSVD(). TIFF0007693133000026.tif26166

[0039] Above, the method of data update when the number of data increases by 1 from N d to N d +1 has been clarified based on FIG. 3, but the disclosed technology is not limited to this. The data processing apparatus 100 according to the disclosed technology may, for example, divide data into a plurality of batches in advance and update the data in batch units. The method of updating data in batch units has the same concept as the method when increasing data one by one sequentially. That is, even when updating data in batch units, the SVD of the Gram matrix may be updated using the first IncrSVD(), and the SVD of the covariance matrix (Σ) may be obtained from the SVD of the correlation matrix using the second IncrSVD().

[0040] (Mahalanobis distance using rank-constrained generalized inverse matrix) As shown in Equation (6), the covariance matrix (Σ) is a positive semi-definite matrix. When the covariance matrix (Σ) is not full rank, the inverse matrix (Σ -1) does not exist. Even when the covariance matrix (Σ) is full rank in numerical calculations, if any of the singular values of the covariance matrix (Σ) is close to 0, so-called division by zero occurs numerically. Also, the problem of seemingly becoming full rank due to numerical errors is as shown in Introduction Part 3. In preparation for such a case, the Mahalanobis distance may be obtained by reducing the dimension of the feature space within the range of the allowable error.

[0041] The data processing apparatus 100 according to the disclosed technology intentionally reduces the rank of the covariance matrix (Σ) to k (k ≤ r). When the number of data is N d , the covariance matrix with its rank reduced to k is given by the following formula. TIFF0007693133000027.tif79166 The “(·) k ” appearing in Equation (27) represents an approximation with the rank reduced to k. At this time, the error due to the rank reduction of the covariance matrix (Σ) can be evaluated as follows. TIFF0007693133000028.tif14166 Finally, the Mahalanobis distance when the rank is k (k ≤ r) is given as follows. TIFF0007693133000029.tif17166

[0042] As described above, the data processing method according to the disclosed technology can benefit from the rank reduction approximation that can reduce the number of parameters to be memorized because of its high affinity with the rank reduction approximation based on singular value decomposition in the learning phase. However, when obtaining the Mahalanobis distance of the subspace reduced in dimension in the inference phase, it is important to know by what projection matrix the subspace is projected and what dimension the space is. The reason why this is important will become clear from the explanation of the “meaningful zero singular value” described in Embodiment 5. The data processing apparatus 100 according to the disclosed technology outputs the following data in a set with the calculated Mahalanobis distance. TIFF0007693133000030.tif12166 That is, the data processing apparatus 100 according to the disclosed technology outputs the calculated Mahalanobis distance, as well as the eigenvectors and eigenvalues related to the covariance matrix (Σ).

[0043] FIG. 4 is a first flowchart showing the processing steps in the parameter update of the data processing apparatus 100 according to Embodiment 1. As shown in FIG. 4, the processing steps of the data processing apparatus 100 according to Embodiment 1 can be divided into ST01 to ST07.

[0044] The processing step (ST01) described as "acquisition of new image feature amount (x)" is a processing step performed via the input interface 110. In ST01, the data processing apparatus 100 acquires, via the input interface 110, for example, the (N + 1)-th data x. d +1st data x Nd+1 to acquire.

[0045] The processing step (ST02) described as "update of the number of data (N)" is a processing step performed by the processing circuit 120 or the processor 122. In ST02, the processing circuit 120 or the processor 122 updates the count of the number of data from, for example, N to N + 1. d )'s update" and described processing step (ST02) is a processing step implemented by the processing circuit 120 or the processor 122. In ST02, the processing circuit 120 or the processor 122 updates the count of the number of data from, for example, N d from N d +1.

[0046] The processing step (ST03) described as "update of the average feature amount (μ)" is a processing step performed by the processing circuit 120 or the processor 122. In ST03, the processing circuit 120 or the processor 122 performs the "update of the average feature amount" shown in Equation (25).

[0047] The processing step (ST04) described as "Update of U, S, V for the Gram matrix" is a processing step performed by the processing circuit 120 or the processor 122. In ST04, the processing circuit 120 or the processor 122 updates the SVD for the Gram matrix using the first IncrSVD() shown in Equation (23).

[0048] "Calculation of U, S for Σ" 2 The processing step (ST05) described as "Calculation of U, S for Σ" is a processing step performed by the processing circuit 120 or the processor 122. In ST05, the processing circuit 120 or the processor 122 calculates the SVD for the covariance matrix (Σ) when the number of data is N d +1 using the second IncrSVD() shown in Equation (26).

[0049] "Σ" ―1 "Storage of U, S for Σ" -2 The processing step (ST06) described as "Storage of U, S for Σ" is a processing step performed using the memory 124 or an external storage device. In ST06, the data processing apparatus 100 stores the eigenvectors (U) and eigenvalues (S ―1 ) for the latest Σ in the memory 124 or an external storage device. -2

[0050] FIG. 5 is a second flowchart showing the processing steps in the parameter update of the data processing apparatus 100 according to Embodiment 1. The processing steps shown in FIG. 5 are the same as the processing steps shown in FIG. 4. However, as shown in FIG. 5, the processing steps shown in FIG. 5 use the k-dimensional feature space mapped by the fixed (U) k T . The image feature amounts (with a bar accent symbol attached to x) handled by the flowchart shown in FIG. 5 are given by the following equation. TIFF0007693133000031.tif28166 However, the (U) appearing in Equation (31) k Tis an empirically obtained matrix, N f performs a mapping from an N-dimensional space to a k-dimensional space. (U) k T When the data is rich enough, that is, when it contains sufficient information to reproduce the properties of the class to which the data belongs, it may be a matrix composed of the first to k-th right singular vectors with respect to the covariance matrix (Σ) (see Equation (27)). A method for determining whether the data is rich enough will be clarified by the description given in Embodiment 6.

[0051] FIG. 6 is a flowchart showing the processing steps in the distance calculation of the data processing apparatus 100 according to Embodiment 1. FIGS. 4 and 5 represent the processing steps performed in the learning phase of the data processing apparatus 100, while FIG. 6 represents the processing steps performed in the inference phase of the data processing apparatus 100. As shown in FIG. 6, the processing steps in the inference phase can be divided into ST11 to ST13. Although not particularly shown, the data processing apparatus 100 in the learning phase and the data processing apparatus 100 in the inference phase may be different apparatuses. That is, there may be a data processing apparatus 100 that performs inference separately from the data processing apparatus 100 that performs learning.

[0052] "Acquisition of sample (x target )" described in the processing step (ST11) is a processing step performed by the data processing apparatus 100 that performs inference. The data processing apparatus 100 that performs inference acquires, in ST11, x target which is the image feature amount for which the Mahalanobis distance is to be obtained.

[0053] "Loading of U and S -1 for Σ -2 " described in the processing step (ST12) is a processing step performed by the data processing apparatus 100 that performs inference. The data processing apparatus 100 that performs inference loads, in ST12, Σ -1For U,S -2 Read it in.

[0054] The processing step (ST13) described as "calculation of Mahalanobis distance (d M )" is a processing step performed by the data processing apparatus 100 that performs inference. The data processing apparatus 100 that performs inference calculates the Mahalanobis distance (d M ), or the square of the Mahalanobis distance (d M 2 ) (see equations (7), (21), (29)).

[0055] In order to clarify the feature space in which the calculated Mahalanobis distance is defined, the data processing apparatus 100 that performs inference outputs the data shown in equation (30), that is, (U) k and (S 2 ) k or (S ―2 ) k along with the calculated Mahalanobis distance.

[0056] (Simple numerical example) The data processing method according to the disclosed technology will be further clarified by the following simple numerical example. Suppose that the data belonging to a certain class is given as follows. TIFF0007693133000032.tif21166 The numerical example shown in equation (32) is for the case where the number of data (N d ) increases from 3 to 4. Suppose that the average feature amount and the SVD of the Gram matrix when the number of data is 3 are given as follows. TIFF0007693133000033.tif55166 Note that due to the limitation of the paper, the numerical examples in this specification are shown up to 7 digits.

[0057] The SVD of the Gram matrix is updated by the first IncrSVD(). TIFF0007693133000034.tif40166 As shown in Equation (34), due to the symmetry of the Gram matrix, the left singular vector and the right singular vector are identical. The SVD obtained by the first IncrSVD() is the SVD of the Gram matrix when the number of data is 4. TIFF0007693133000035.tif43166

[0058] The average feature amount is updated as follows. TIFF0007693133000036.tif20166

[0059] By the second IncrSVD(), the SVD of the covariance matrix (Σ) is obtained as follows. TIFF0007693133000037.tif42166 The SVD obtained by the second IncrSVD() is the SVD of the covariance matrix (Σ) when the number of data is 4. TIFF0007693133000038.tif50166

[0060] The technical feature of the data processing apparatus and the data processing method according to Embodiment 1 is that the algorithm for updating the covariance matrix (Σ) in the learning phase has an affinity with singular value decomposition. More specifically, the data processing apparatus and the data processing method according to Embodiment 1 include a function called IncrSVD() that has an affinity with singular value decomposition in the algorithm for updating the covariance matrix (Σ) in the learning phase. By having this technical feature, the data processing apparatus and the data processing method according to Embodiment 1 can benefit from singular value decomposition in that they can address the problem of seemingly becoming full rank due to numerical errors (see Introduction Part 3).

[0061] The technical feature of the data processing apparatus and the data processing method according to Embodiment 1 is that the algorithm for updating the scatter-covariance matrix (Σ) in the learning phase has an affinity with the low-rank approximation based on singular value decomposition. More specifically, the data processing apparatus and the data processing method according to Embodiment 1 include a function called IncrSVD() in the algorithm for updating the scatter-covariance matrix (Σ) in the learning phase, which has an affinity with the low-rank approximation based on singular value decomposition. By having this technical feature, the data processing apparatus and the data processing method according to Embodiment 1 can benefit from the low-rank approximation based on singular value decomposition, which can slim down the information to the required amount in the learning phase.

[0062] The data processing apparatus and the data processing method according to Embodiment 1 are applied to anomaly detection using PaDiM, and their effects have been verified. In this application example, the image is divided into small regions of 56×56. For each small region, the scatter-covariance matrix (Σ) is updated. As described in Non-Patent Document 1 regarding PaDiM, Wide ResNet50 is used in the CNN, and the generated image feature amount has a feature amount length (N f ) of 1792. Even simply calculating, the memory required for updating in the learning phase is 56×56×1792×1792 = 40 [GB]. Considering that the scatter-covariance matrix (Σ) is a symmetric matrix, about 20 [GB] of memory is still required. In this application example, the data processing apparatus and the data processing method according to Embodiment 1 have demonstrated that the dimension of the update in the learning phase can be reduced to k = 20.

[0063] Embodiment 2. The data processing apparatus and the data processing method according to Embodiment 2 are modifications of the data processing apparatus and the data processing method according to the present disclosure technology. Unless otherwise specified, the same symbols as those used in Embodiment 1 are used in Embodiment 2. Also, in Embodiment 2, the descriptions overlapping with those in Embodiment 1 are omitted as appropriate. To distinguish from the method shown in Embodiment 1, in this specification, the data processing method according to Embodiment 2 shall be referred to as "IncrPCA". FIG. 7 is an explanatory diagram showing the data processing method according to Embodiment 2.

[0064] By the way, the core idea of IncrSVD() is to introduce the following extended system. TIFF0007693133000039.tif13166 The method of introducing the extended system is also effective for the cases handled by the present disclosed technology. When the number of data is N d and it is assumed that the SVD of the Gram matrix is given as in Equation (22). However, considering that the Gram matrix is a positive semi - definite matrix represented by a quadratic form, the SVD of the Gram matrix can be expressed as follows. TIFF0007693133000040.tif15166 Note that the SVD of the Gram matrix may be approximated by rank reduction. The feasible conditions for performing rank reduction approximation will be clarified in Embodiment 5. N d When the (N + 1)-th data is obtained, the Gram matrix can be represented by the following extended system. TIFF0007693133000041.tif22166 Equation (41) suggests the following. TIFF0007693133000042.tif14166 Here, the script font "SVD" appearing in Equation (42) represents a function for obtaining the singular value decomposition.

[0065] As described above, when the number of data is N d to N dRegarding the data update method when it increases by 1 to +1, an expansion system was introduced and clarified, but the disclosed technology is not limited to this. The data processing apparatus 100 according to the disclosed technology may, for example, divide data into a plurality of batches in advance, add data in batch units, and create an expansion system. The method of introducing an expansion system in batch units has the same concept as the method when increasing data one by one sequentially. Batch processing means that when the number of data is N d and n b new data are added all at once. When performing such data processing, Equation (42) can be rewritten as follows. TIFF0007693133000043.tif27166 Here, I(n b ) that appears in Equation (43) is an identity matrix of size n b ×n b . As described above, the SVD of the Gram matrix may be approximated to a low rank in k (k < r) dimensions. The fact that it is approximated to a low rank in k dimensions means that only the top k of the singular values of the Gram matrix arranged in descending order are used (see, for example, Equation (27)). However, since the update of the Gram matrix becomes a loop process, the error caused by the low rank approximation accumulates. The inventor of the disclosed technology reduced the original dimension of N f = 1792 to k = 20 and updated the Gram matrix by loop processing in the anomaly detection process by PaDiM. The inventor of the disclosed technology found that in this application example, this cumulative error is within an acceptable range.

[0066] (Simple numerical example) By using the same numerical example as that shown in Embodiment 1, the data processing method according to Embodiment 2 becomes clearer. When the number of data is 3, the U of the Gram matrix G0 and S G0 are given as follows. TIFF0007693133000044.tif24166 Note that U G0 is the same as U0.

[0067] Applying numerical examples, Equation (42) is calculated as follows. TIFF0007693133000045.tif30166 Furthermore, the matrix calculated by Equation (45) can be expressed in the following SVD form. TIFF0007693133000046.tif76166 Here, the subscript "3" in the lower right that appears in Equation (46) is simply a number for identifying which matrix the SVD is for. The first and second singular vectors of U3 calculated by Equation (46) are the same as the first and second singular vectors of U1 calculated by Equation (34). Furthermore, when the first and second singular values of S3 calculated by Equation (46) are squared respectively, they match the first and second singular values of S1 calculated by Equation (34). TIFF0007693133000047.tif34166 As shown in Equation (47), the data processing device according to Embodiment 2 expands the dimension of the matrix to (N f +n b )×(N f +n b ). However, in the SVD of the Gram matrix, n b singular values become zero. Therefore, finally, a decomposition form of size N f ×N f is obtained. The processing steps in the data processing method after Equation (47) are the same as those in the data processing method according to Embodiment 1.

[0068] The technical features peculiar to the data processing device and data processing method according to Embodiment 2 are that in the learning phase, an algorithm of introducing an extended system to update the Gram matrix is provided. Due to this technical feature, the data processing apparatus and the data processing method according to Embodiment 2 have, in addition to the effects described in Embodiment 1, the effect that it can be realized by using a general-purpose SVD function instead of IncrSVD().

[0069] Embodiment 3 The data processing apparatus and the data processing method according to Embodiment 3 are modifications of the data processing apparatus and the data processing method according to the present disclosed technology. Unless otherwise specified, in Embodiment 3, the same reference numerals as those used in the previous embodiments are used. Also, in Embodiment 3, descriptions overlapping with the previous embodiments are appropriately omitted. To distinguish from the methods shown in the previous embodiments, in this specification, the data processing method according to Embodiment 3 shall be referred to as "GPU-oriented IncrPCA".

[0070] The data processing apparatus according to the present disclosed technology may use a GPU (Graphics Processing Unit) instead of a CPU (Central Processing Unit). The advantage of using a GPU lies in that a function for obtaining an SVD (hereinafter simply referred to as an "SVD function") is already prepared as a library (for example, the cuSOLVER library of NVidia). The GPU-based SVD function is fast if the size of the matrix is 32×32 or less.

[0071] The scenario assumed in Embodiment 3 is, for example, a scenario where even if a low-rank approximation is performed with k = 32 or less, the error is small enough to be negligible. However, for the sake of easy understanding of the data processing method according to Embodiment 3, in this specification, first, the mathematical formula without performing a low-rank approximation is described.

[0072] (Prerequisite knowledge of the data processing method according to Embodiment 3) The singular value decomposition of an arbitrary matrix (Z) is Z T the eigenvalue decomposition of Z and the eigenvalue decomposition of ZZ T can be obtained by. The form of the singular value decomposition of Z is USVT Assume that it is so. Z T Z can be transformed as follows. TIFF0007693133000048.tif25166 Here, if the i-th column of V is denoted as v i then, from Equation (48), the following relationship can be obtained. TIFF0007693133000049.tif10166 That is, v i and σ i 2 are respectively the eigenvector and eigenvalue of Z T Z. Similarly, ZZ T can be transformed as follows. TIFF0007693133000050.tif26166 Here, if the i-th column of U is denoted as ui, then from Equation (50), the following relationship can be obtained. TIFF0007693133000051.tif11166 As described above, the singular value decomposition of an arbitrary matrix (Z) can be obtained by T the eigenvalue decomposition of Z and T the eigenvalue decomposition of ZZ.

[0073] For example, assume that the size of Z is p×q and it is a horizontally long matrix, that is, q > p. In this case, Z T The size of Z is small, p×p, but T the size of ZZ is large, q×q. In such a case, T it is advisable to perform only the eigenvalue decomposition of Z to calculate the matrix (V) related to the right singular vector and the matrix (S) related to the singular values. The matrix (U) related to the left singular vector may be calculated by the following calculation. TIFF0007693133000052.tif24166 Here, if a low-rank approximation is performed in the k (k < r) dimension, all the assumed k singular values become non-zero, and S -1can be made to always exist. Note that Z T In the eigenvalue decomposition of Z, if there are no non-zero eigenvalues, S can be calculated without performing a low-rank approximation -1 .

[0074] (Details of the data processing method according to Embodiment 3) FIG. 8 is an explanatory diagram showing the data processing method according to Embodiment 3. The data processing apparatus according to Embodiment 3 defines the following matrix that appears in Equation (43) as M. TIFF0007693133000053.tif39166 Note that M is a matrix (X) consisting of data up to d +n b . The subscript b of X is a sequential number assigned for each batch. The number of data in the b-th batch is n b .

[0075] Assume that M can be SVD as follows. TIFF0007693133000054.tif17166 However, at this stage, it is assumed that the SVD form has not yet been obtained. Note that the subscript "G1" that appears in Equation (54) is simply a symbol for identifying which matrix the SVD is for.

[0076] The matrix (V G1 ) related to the right singular vectors of M and the matrix (S G1 ) related to the singular values can be obtained by eigenvalue decomposition of M as shown in Equations (48) and (49). T Also, the matrix (U G1 ) related to the left singular vectors of M is calculated from the following relational expression. TIFF0007693133000055.tif10166 Note that if all the eigenvalues of M T are non-zero, the matrix (S G1 ​) has an inverse matrix.

[0077] The data processing apparatus according to Embodiment 3 updates M each time the number of data increases. The updated M is given as follows. TIFF0007693133000056.tif39166 In Equation (56), the number of data in the (b + 1)-th batch is n b+1 is. The update of M shown in Equation (56) corresponds to the update of the Gram matrix using the first IncrSVD() in Embodiment 1. When performing singular value decomposition of M, low-rank approximation may be used. However, since the update of M becomes a loop process, the errors caused by low-rank approximation accumulate. The inventor of the present disclosure technology, in the anomaly detection process by PaDiM, originally had N f = 1792 dimensions were reduced to k = 20, and the update of M was performed by loop processing. The inventor of the present disclosure technology found that in this application example, this cumulative error is within an acceptable range.

[0078] After reflecting a sufficient number of data, the operation of obtaining the SVD of the covariance matrix (Σ) from the SVD of the Gram matrix is also referred to as Finalization. At the Finalization stage, assume that the SVD of the Gram matrix is obtained as follows. TIFF0007693133000057.tif15166 The subscript “G” in the lower right in Equation (57) is a symbol to emphasize that it is the SVD of the Gram matrix. Also, as shown in Equation (57), the SVD of the Gram matrix is approximated to be reduced to k dimensions. N all is the total number of data at the time of Finalization. In Finalization, the SVD of the covariance matrix (Σ) is given as follows using the SVD function. TIFF0007693133000058.tif21166 However, the SVD shown in Equation (58) also includes errors due to dimensionality reduction approximation. Also, the average feature amount (μ) appearing on the right side of Equation (58) is reduced to k dimensions considering the matrix size in matrix addition and subtraction. Equation (58) can also be said to give the SVD of the covariance matrix (Σ) when the feature space is reduced to k dimensions. This SVD function may be a GPU-based SVD function.

[0079] The technical features specific to the data processing apparatus and the data processing method according to Embodiment 3 are T calculating the singular value decomposition of M based on the eigenvalue decomposition of M. Due to this technical feature, the data processing apparatus and the data processing method according to Embodiment 3 also have the effect that, in addition to the effects described in the existing embodiments, they can be realized by a function for obtaining a general eigenvalue decomposition.

[0080] Embodiment 4. The data processing apparatus and the data processing method according to Embodiment 4 are modified examples of the data processing apparatus and the data processing method according to the present disclosure technology. Unless otherwise specified, the same reference numerals as those used in the existing embodiments are used in Embodiment 4. Also, in Embodiment 4, descriptions overlapping with the existing embodiments are appropriately omitted.

[0081] (Verification function) The data processing apparatus and the data processing method according to the present disclosure technology may include a verification function. The data processing apparatus and the data processing method according to the present disclosure technology may obtain a correlation matrix for verification by the following sequential method that is not in the SVD format. TIFF0007693133000059.tif40166 The sequential update formula of the correlation matrix shown in Equation (59) has the same form as the sequential update formula of the average feature amount (μ) shown in Equation (25). Using the expectation function E(), Equation (25) is expressed as follows. TIFF0007693133000060.tif15166

[0082] Using equations (59) and (60), the covariance matrix (Σ) for verification can be obtained by the following sequential method that is not in SVD form. TIFF0007693133000061.tif37166 The updated covariance matrix (Σ) obtained by equation (61) may be used to verify whether the SVD form obtained by the data processing method shown in the existing embodiment was correctly obtained.

[0083] The technical features specific to the data processing apparatus and the data processing method according to Embodiment 4 are that the correlation matrix for verification can be obtained directly and sequentially. Here, the meaning of "directly" is "not in SVD form". Due to this technical feature, in addition to the effects described in the existing embodiments, the data processing apparatus and the data processing method according to Embodiment 4 also have the effect of being able to verify the correctness of the covariance matrix (Σ) calculated in SVD form in the learning phase.

[0084] Embodiment 5. The data processing apparatus and the data processing method according to Embodiment 5 are modified examples of the data processing apparatus and the data processing method according to the present disclosure technology. Unless otherwise specified, the same reference numerals as those used in the existing embodiments are used in Embodiment 5. Also, in Embodiment 5, descriptions overlapping with the existing embodiments are appropriately omitted.

[0085] When the covariance matrix (Σ) after Finalization is full rank, the squared Mahalanobis distance (d target ) for the target sample (x M 2 ) is expressed as follows using the generalized inverse matrix (Σ - ) (see also equation (21)). TIFF0007693133000062.tif78166 Here, the magnitudes of the singular values are σ1 ≧ … ≧ σ Nf That is. In this way, the Mahalanobis distance is obtained by the generalized inverse matrix (Σ - ) of the covariance matrix (Σ), but the generalized inverse matrix (Σ - ) has the smallest first singular value (1 / σ1 2 ) and the largest final singular value (1 / σ Nf 2 ) is the largest. The fact that the final singular value (1 / σ Nf 2 ) is the largest means that the term (component) for the final singular value (1 / σ Nf 2 ) has the greatest influence on the Mahalanobis distance.

[0086] In the learning phase, it can also be said that the covariance matrix (Σ) representing the properties of a specific class is important in the order of the magnitudes of the singular values, that is, in the order of σ1, …, σ Nf . Therefore, in the learning phase, there is no problem if the assumed feature space is a k-dimensional subspace. On the other hand, in the inference phase, it is important to set the feature space as a full N f -dimensional space.

[0087] (meaningful zero singular value) Even if the dimension of the space spanned by all the deviation vectors (y1, …, y Nall ) belonging to the training data is r (r < N f ), the deviation vector (y target ) for the target does not necessarily belong to this r-dimensional space. For example, when considering a class consisting of images in a normal state for a certain object, assume that the dimension of the space spanned by the deviation vectors (y1, …, y Nall ) for the images in the normal state is r (r < N f ). Even so, the deviation vector (y target) does not necessarily belong to this r-dimensional space. In such a case, the zero singular value of the scatter-covariance matrix (Σ) for the class consisting of normal-state images is a "meaningful zero singular value". On the contrary, for any deviation vector (y target ) with respect to, if it always belongs to this r-dimensional space, dimensions larger than r are redundant, and the zero singular value of the scatter-covariance matrix (Σ) is a meaningless zero singular value.

[0088] Assume that the rank of the scatter-covariance matrix (Σ) for the class consisting of normal-state images is r (r < N f ). In this case, the data processing apparatus and data processing method in the inference phase may perform the following determination process as a method for dealing with "meaningful zero singular values" before calculating the Mahalanobis distance (d M ). TIFF0007693133000063.tif41166 However, ε (epsilon) appearing in Equation (63) is a threshold value such as machine epsilon. When the condition shown in Equation (63) is true, the deviation vector (y target ) for the target does not belong to the class consisting of normal-state images. If it is clear that the target does not belong to the class, there is no need to obtain the Mahalanobis distance (d M ) intentionally. Incidentally, the Mahalanobis distance in this case, strictly speaking, becomes ∞ (infinity) as a result of division by the zero singular value. When the condition shown in Equation (63) is true, the zero singular value of the scatter-covariance matrix (Σ) for the class consisting of normal-state images is a non-truncatable singular value, that is, a "meaningful zero singular value". Only when the condition shown in Equation (63) is false, the data processing apparatus and data processing method in the inference phase execute the process of calculating the Mahalanobis distance (d M ) defined in the r-dimensional feature space.

[0089] When performing a low-rank approximation that regards the dimension as k (k < r), the data processing apparatus and data processing method according to the disclosed technology apply Equation (63) to perform the following determination process. TIFF0007693133000064.tif40166 When a low-rank approximation regarding the dimension as k is feasible, the approximate value of the Mahalanobis distance is given by the following equation based on the rank-constrained generalized inverse matrix. TIFF0007693133000065.tif46166 The variable with a bar accent mark attached to y defined in the second equation of Equation (65) shall be referred to as the "normalized deviation vector for the target" or simply the "normalized deviation vector" in this specification. Note that even when no low-rank approximation is performed, the same vector shall be referred to as the "normalized deviation vector" (see Equation (62)).

[0090] Equation (64) gives the conditions under which a low-rank approximation may be performed when calculating the Mahalanobis distance. The conditions under which a low-rank approximation may be performed are as follows. TIFF0007693133000066.tif22166 Equation (66) represents that when looking at the coordinates of y in the N-dimensional space defined by the basis vectors from u1 to u Nf to u f , the absolute values of the coordinate components from the (k + 1)-th to the N target -th are all less than ε (epsilon). That is, the condition under which a low-rank approximation may be performed is that the deviation vector (y f ) for all targets is an element of the k-dimensional subspace spanned by the deviation vectors (y1,..., y target ). Therefore, "the deviation vector (y Nall ) for all targets to be measured in the future is an element of the k-dimensional subspace spanned by the deviation vectors (y1,..., y target )" NallUnless there are circumstances such as being able to empirically predict that it "becomes an element of a k-dimensional subspace spanned by ()", low-rank approximation should not be performed. In this specification, the condition given by Equation (66) shall be referred to as the "feasible condition for performing low-rank approximation". When it is not known whether the target (application example) to which the present disclosure technology is to be applied satisfies the feasible condition for performing low-rank approximation, in the inference phase, the data processing apparatus and the data processing method may perform the determination process shown in Equation (63) or Equation (64).

[0091] The inventors of the present disclosure technology found that in the anomaly detection process by PaDiM, even when reducing the original dimension of N f = 1792 to k = 20, the "feasible condition for performing low-rank approximation" shown in Equation (66) is satisfied.

[0092] The technical feature peculiar to the data processing apparatus and the data processing method according to Embodiment 5 is that it includes a method for dealing with "meaningful zero singular values". Due to this technical feature, the data processing apparatus and the data processing method according to Embodiment 5 can solve learning problems such as "classification" and "clustering" using a full N f dimensional space as the feature space even when the scatter covariance matrix (Σ) related to a certain class is not full rank in the inference phase.

[0093] Embodiment 6. The data processing apparatus and the data processing method according to Embodiment 6 are modified examples of the data processing apparatus and the data processing method according to the present disclosure technology. Unless otherwise specified, in Embodiment 6, the same reference numerals as those used in the previous embodiments are used. Also, in Embodiment 6, descriptions overlapping with the previous embodiments are appropriately omitted.

[0094] By the way, in FIG. 5 according to Embodiment 1, the fixed eigenvector ((U) k TThe flowchart in the case of using the k-dimensional feature space mapped by ( ) was shown. As described above, in order to use the fixed k-dimensional subspace projected by the fixed eigenvector, it was necessary that the data was rich enough, that is, the data had acquired sufficient information to reproduce the properties of the class to which the data belonged. The data processing apparatus and data processing method according to the disclosed technology include means for determining whether the data is rich enough, specifically, an end condition for the loop process related to data update.

[0095] (End condition for the loop process related to data update) When the number of data is N d +1, based on the information when the number of data is N d The end condition for the loop process related to data update is given, for example, as follows. TIFF0007693133000067.tif28166 Here, ε appearing in Equation (67) σ is a threshold for singular values, and ε u is a threshold for eigenvectors. ε σ and ε u may be the same value or different values. Also, the singular values (σ i ) and eigenvectors (u i ) appearing in Equation (67) are related to the following scatter covariance matrix (Σ). TIFF0007693133000068.tif43166

[0096] The determination of whether the data is rich enough may indirectly observe the singular values and eigenvectors of the scatter covariance matrix (Σ) even if it is not possible to directly observe the singular values and eigenvectors of the scatter covariance matrix (Σ). A method for indirectly observing the singular values and eigenvectors of the scatter covariance matrix (Σ) is specifically to observe the correlation matrix.

[0097] In the case of the data processing apparatus and the data processing method according to Embodiment 1, for the end condition represented by Equation (67), it is preferable to use the singular values and singular vectors for the SVD of the correlation matrix derived from the SVD of the Gram matrix represented by Equation (23).

[0098] In the case of the data processing apparatus and the data processing method according to Embodiment 2, for the end condition represented by Equation (67), it is preferable to use the singular values and singular vectors for the SVD of the correlation matrix derived from the SVD of the Gram matrix represented by Equation (42) or Equation (43).

[0099] In the case of the data processing apparatus and the data processing method according to Embodiment 3, for the end condition represented by Equation (67), M T It is preferable to use the singular values and singular vectors for the SVD of the correlation matrix derived from the SVD of the Gram matrix obtained by eigenvalue decomposition of M or the like.

[0100] The technical feature peculiar to the data processing apparatus and the data processing method according to Embodiment 6 is that it has an end condition for the loop process related to data update. Due to this technical feature, the data processing apparatus and the data processing method according to Embodiment 6 also have the effect that they can determine the end of the loop process based on the determination of whether the data is sufficiently rich.

Industrial Applicability

[0101] The data processing apparatus and the data processing method according to the present disclosure technology can be applied to a defect inspection apparatus that performs abnormality determination, for example, a defect inspection apparatus for a photomask for semiconductors, and have industrial applicability.

Explanation of Signs

[0102] 100 Data processing apparatus, 110 Input interface, 120 Processing circuit, 122 Processor, 124 Memory, 130 Output interface.

Claims

1. A data processing device comprising a processing circuit, wherein the processing circuit sequentially updates a Gram matrix in the form of SVD in a learning phase, and the processing circuit calculates a covariance matrix in the form of SVD based on the SVD of the Gram matrix in the Finalization of the learning phase.

2. A data processing device comprising a processor that executes a program, wherein the processor sequentially updates a Gram matrix in the form of SVD in a learning phase by executing the program, and the processor calculates a covariance matrix in the form of SVD based on the SVD of the Gram matrix in the Finalization of the learning phase by executing the program.

3. The program includes a function that takes the SVD of Z, A, and B as inputs and outputs the SVD of Z + AB T provided that A, B, and Z are matrices respectively. The data processing device according to claim 2.

4. The program includes a function that outputs the SVD of an extended system. The data processing device according to claim 2.

5. The program includes a function that outputs the eigenvalue decomposition of a matrix (M T M) represented in a quadratic form. The data processing device according to claim 2.

6. The program includes a function that calculates a correlation matrix for verification not in the form of SVD and sequentially. The data processing device according to any one of claims 2 to 5.

7. The program includes a function for dealing with zero singular values that have meaning. The data processing apparatus according to any one of claims 2 to 5.

8. The program includes a function for determining the end of loop processing based on a determination of whether learning data is sufficiently rich. The data processing apparatus according to any one of claims 2 to 5.

9. A data processing apparatus, sequentially updates a Gram matrix in the form of SVD in a learning phase, and calculates a scatter covariance matrix in the form of SVD based on the SVD of the Gram matrix in the Finalization of the learning phase. A data processing method.

10. The data processing apparatus takes the SVD of Z, A, and B as inputs, and includes a numerical calculation that outputs the SVD of Z + AB T where A, B, and Z are each matrices. The data processing method according to claim 9. The data processing method according to claim 9.

11. The data processing apparatus includes a numerical calculation that outputs the SVD of an extended system. The data processing method according to claim 9.

12. The data processing apparatus includes a numerical calculation that outputs an eigenvalue decomposition of a matrix (M T M) represented in a quadratic form. The data processing method according to claim 9.

13. The data processing apparatus further includes a numerical calculation for sequentially calculating a correlation matrix for verification, not in the form of SVD. The data processing method according to any one of claims 9 to 12.

14. The data processing apparatus includes a process for dealing with zero singular values that have meaning. The data processing method according to any one of claims 9 to 12.

15. The data processing apparatus includes a process of determining the end of the loop process based on a determination of whether the learning data is sufficiently rich. The data processing method according to any one of claims 9 to 12.

Citation Information

Patent Citations

  • Method for incremental singular value decomposition

    JP2003316764A

  • Information processing device and microscope system

    JP2020144109A