Data processing device and data processing method

WO2026196618A1PCT designated stage Publication Date: 2026-09-24MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/022239
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2025-06-20
Publication Date
2026-09-24

Smart Images

  • Figure JP2025022239_24092026_PF_FP_ABST
    Figure JP2025022239_24092026_PF_FP_ABST
Patent Text Reader

Abstract

A data processing device (1) comprises a general-purpose processor and a neural network processor, and calculates the singular value decomposition of a prescribed matrix using the one-sided Jacobi algorithm. From the start of the singular value decomposition calculation until a first convergence criterion is satisfied, the neural network processor performs prescribed matrix operations that are carried out iteratively during the process of the singular value decomposition calculation using the one-sided Jacobi algorithm. After the first convergence criterion is satisfied and until a second convergence criterion more stringent than the first convergence criterion is satisfied, the general-purpose processor performs the prescribed matrix operations.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing device and data processing method

[0001] This disclosure relates to a data processing device and a data processing method.

[0002] Visual inspection is an inspection that checks the appearance of a part or product in order to maintain or guarantee the quality of that part or product. In this inspection, the main focus is on checking for cosmetic defects such as foreign matter, dirt, scratches, burrs, chips, or deformations attached to the surface of the part or product, and an abnormality determination is made based on the results. Conventional visual inspection processing algorithms employ a method of determining abnormalities based on the distance (Mahalanobis distance) between the feature quantities obtained from the image of the part or product to be inspected and the feature quantity distribution of a normal image learned by a convolutional neural network (CNN), etc.

[0003] In this process, the features identified by the feature data used to learn the feature distribution of normal images can be high-dimensional. Therefore, to reduce the computational load on the device that performs anomaly detection and to reduce the memory capacity, it is necessary to compress the dimensionality of the high-dimensional features. Conventional devices perform low-rank approximation calculations of features by performing singular value decomposition on feature data expressed in matrix form. In this regard, for example, Patent Document 1 discloses a method in which features are sequentially provided by inputting several input images at a time to the device, and singular value decomposition calculations are performed sequentially. Furthermore, in the device of Patent Document 1, the singular value decomposition calculation is performed by effectively utilizing the computing resources of a GPU (Graphics Processing Unit) to speed up processing.

[0004] International Publication No. 2024-166339

[0005] In visual inspection processes, the calculation of singular value decomposition (SDL) is particularly computationally intensive, and therefore, large computing devices have traditionally been used. However, there is a demand for a more convenient way to perform visual inspection processes, including SDL calculation, using a smaller computing device. To meet this demand, it is necessary to achieve both power saving and miniaturization of the device, and the performance and accuracy of the SDL calculation process required to meet on-site inspection standards. However, achieving both of these is difficult with a small computing device. For example, the device described in Patent Document 1 uses a GPU, so processing is fast, but power consumption is high, and miniaturization is difficult because a cooling fan is required.

[0006] This disclosure was made to solve the above-mentioned problems, and aims to achieve both power saving and miniaturization of the device that performs singular value decomposition calculations, and the performance and accuracy of said calculations.

[0007] The data processing device according to this disclosure comprises a general-purpose processor and a neural network processor, and is a data processing device that performs singular value decomposition calculation on a predetermined matrix using the one-sided Jacobi method, characterized in that, from the start of the singular value decomposition calculation until a first convergence criterion is met, the neural network processor performs predetermined matrix operations that are iteratively performed in the calculation process of singular value decomposition using the one-sided Jacobi method, and from the time the first convergence criterion is met until a second convergence criterion, which has stricter convergence criterion conditions than the first convergence criterion, is met, the general-purpose processor performs the predetermined matrix operations.

[0008] According to this disclosure, the above configuration makes it possible to achieve both power saving and miniaturization of the device that performs singular value decomposition calculations, and the performance and accuracy of said calculations.

[0009] This figure shows an example of the hardware configuration of the data processing device according to Embodiment 1. This is a block diagram showing an example configuration of a data processing system that includes the data processing device according to Embodiment 1. This is a block diagram showing an example configuration of the low-rank approximation unit and matrix calculation unit in Embodiment 1. This is a flowchart showing an example of the operation of the low-rank approximation unit and matrix calculation unit in Embodiment 1.

[0010] The embodiments will be described in detail below with reference to the drawings. Embodiment 1. The data processing device according to Embodiment 1 comprises a system having a learning function (hereinafter referred to as the "learning system") and a system having an inference function (hereinafter referred to as the "inference system"). When image data (hereinafter referred to as "image data") is input to the learning system, it performs machine learning using the input image data to generate a trained model, and by performing the above machine learning, it generates data representing the image features in matrix form (hereinafter referred to as "feature data"). The learning system outputs the generated trained model to the inference system, performs singular value decomposition on the matrix identified by the generated feature data, and compresses the dimensionality of the features by performing low-rank approximation of the feature data based on the results. The learning system generates a low-rank matrix from the dimensionally compressed features and outputs data representing the generated low-rank matrix to the inference system as a learning result. When image data to be inferred is input to the inference system, it obtains the feature data of the image by inferring the image features from the image data using the trained model output from the learning system. The inference system then calculates the distance between the inferred feature data and the multidimensional space based on the low-rank matrix, based on the feature data obtained through inference and the learning results output from the learning system (data showing the low-rank matrix), and performs anomaly detection based on the calculated result.

[0011] The data processing device 1, which includes the learning system and inference system described above, comprises a general-purpose processor and a neural network processor as hardware resources. In the learning system, as described above, singular value decomposition of the matrix identified by the feature data is performed with the aim of performing low-rank approximation of the feature data. In the data processing device 1, this singular value decomposition calculation is performed in cooperation between the general-purpose processor and the neural network processor. In other words, the data processing device 1 performs the singular value decomposition calculation in a hybrid manner using the two processors described above.

[0012] Figure 1 shows an example of the hardware configuration of the data processing device 1 according to Embodiment 1. As shown in Figure 1, the data processing device 1 according to Embodiment 1 comprises a processor 2, a memory 3, an input interface 4, and an output interface 5 as its hardware configuration.

[0013] Processor 2 is also commonly referred to as a CPU (Central Processing Unit), central processing unit, processing unit, arithmetic unit, microprocessor, or microcomputer. This processor 2 is a general-purpose processor that performs data calculations and controls processing within the data processing unit 1. The data processing unit 1 also includes a neural network processor (shown as NPU 50 in the diagram) in addition to such a general-purpose processor (shown as CPU 10 in the diagram) as processor 2. Unlike the general-purpose processor or GPU, the neural network processor is specialized for matrix operations and AI functions, and is a processor that can perform matrix operations and AI functions at high speed and with low power consumption. Furthermore, the neural network processor is a low-power and efficient processor due to design innovations in terms of dedicated architecture, hardware acceleration, data flow optimization, low-voltage operation, and power management functions. A processor for a neural network may consist of, for example, an NPU, a TPU (Tensor Processing Unit), a DSP (Digital Signal Processor), and other dedicated AI chips.

[0014] Furthermore, GPUs and neural network processors differ in their design purpose and architecture. In terms of design purpose, GPUs are primarily used for general-purpose parallel processing tasks such as graphics rendering, image processing, or scientific computing. Therefore, GPUs are widely used, for example, for deep learning training and inference. On the other hand, neural network processors are specifically designed for AI and machine learning tasks, particularly neural network inference. Neural network processors provide dedicated hardware acceleration and achieve high efficiency in specific AI tasks.

[0015] Furthermore, in terms of architecture, GPUs have numerous cores and a flexible architecture for general-purpose parallel computing. GPUs can perform a wide range of computational tasks using programming models such as CUDA or OpenCL. On the other hand, neural network processors have a dedicated architecture optimized for neural network computations and efficiently perform specific operations such as matrix operations and convolution operations. Therefore, GPUs can be said to be computing devices or processors specialized for real-time image processing, and are more specialized for parallel processing than general-purpose processors, and thus GPUs can also compute matrix operations and AI functions at high speed. However, unlike neural network processors, GPUs generally consume a lot of power and require large cooling fans, making it difficult to miniaturize devices when they are installed. Against this backdrop, the data processing device 1 according to Embodiment 1 employs a neural network processor instead of a GPU as a means to solve the above-mentioned problems.

[0016] Each function of the data processing device 1 is realized by software, firmware, or a combination of both. This software and firmware are written as programs and stored in memory 3. The processor 2 realizes each function by reading and executing the programs stored in memory 3. In other words, the data processing device 1 is equipped with memory 3 for storing programs that, when executed by the processor 2, execute at least each processing step of the flowchart shown in Figure 4, which will be described later. These programs can also be said to cause the processor 2 to execute the procedures and methods of the data processing device 1. Here, memory 3 may be a non-volatile or volatile semiconductor memory such as RAM (Random Access Memory), ROM (Read Only Memory), flash memory, or EPROM (Erasable Programmable Read Only Memory). Furthermore, memory 3 may include disks such as magnetic disks, flexible disks, optical disks, compact disks, minidiscs, and DVDs (Digital Versatile Discs). In addition, memory 3 may be in the form of an HDD (Hard Disk Drive) or an SSD (Solid State Drive).

[0017] The input interface 4 is used when the user of the data processing device 1 performs input operations on the data processing device 1, and is composed of, for example, keywords, a mouse, or a touch panel. The output interface 5 presents the output data from the data processing device 1 to the user of the data processing device 1, and is composed of, for example, a display or a touch panel.

[0018] Next, a data processing system comprising the data processing device 1 will be described. Figure 2 is a block diagram showing an example configuration of a data processing system comprising the data processing device 1. The data processing system mainly consists of a sensor 60 and a data processing device 1 equipped with the general-purpose processor and neural network processor described above. In the following, the combination of the general-purpose processor and the neural network processor will be described as the combination of CPU 10 and NPU 50. As shown in Figure 2, the CPU 10 implements the functions of the learning system 20 and the inference system 30. The NPU 50 implements the functions of the matrix operation unit 51.

[0019] The sensor 60 is composed of an imaging device, such as a camera capable of capturing images, and is connected to the data processing device 1 in a communicative manner. The sensor 60 outputs image data obtained by capturing images to the learning system 20 and inference system 30 of the data processing device 1.

[0020] The learning system 20 includes a learning-side input / output unit 21, a CNN learning unit 22, and a low-rank approximation unit 23. The inference system 30 includes an inference-side input / output unit 31, a CNN inference unit 32, and an anomaly distribution calculation unit 33.

[0021] In the learning system 20, the learning-side input / output unit 21 acquires image data from the sensor 60 and outputs the acquired data to the CNN learning unit 22. The CNN learning unit 22 acquires image data from the learning-side input / output unit 21 and uses the acquired image data to perform training on a convolutional neural network and calculate image features identified by the image data. The CNN learning unit 22 outputs the convolutional neural network obtained through training (hereinafter referred to as the "trained model") to the CNN inference unit 32 of the inference system 30. The CNN learning unit 22 also outputs data indicating the calculated features (feature data) to the low-rank approximation unit 23. The low-rank approximation unit 23 acquires the feature data from the CNN learning unit 22 and outputs the acquired feature data to the matrix calculation unit 51 of the NPU 50. The matrix calculation unit 51 of the NPU 50 acquires the feature data from the low-rank approximation unit 23 of the learning system 20 and obtains a Gram matrix by performing predetermined matrix calculations. The matrix calculation unit 51 outputs data representing the Gram matrix, which is the result of matrix calculation, to the low-rank approximation unit 23. The low-rank approximation unit 23 obtains data representing the Gram matrix from the matrix calculation unit 51 of the NPU 50 and calculates singular value decomposition of the features (matrices) identified by the feature data based on the Gram matrix and feature data identified by the obtained data. The low-rank approximation unit 23 also compresses the dimensionality of the features by performing low-rank approximation of the feature data based on the results of the singular value decomposition calculation. The low-rank approximation unit 23 then generates a low-rank matrix from the dimensionally compressed features and outputs data representing the generated low-rank matrix as a learning result to the anomaly distribution calculation unit 33 of the inference system 30.

[0022] In the inference system 30, the inference-side input / output unit 31 acquires image data from the sensor 60 and outputs the acquired image data to the CNN inference unit 32. The CNN inference unit 32 acquires image data from the inference-side input / output unit 31 and acquires a trained model from the CNN learning unit 22 of the learning system 20. Using the trained model acquired from the CNN learning unit 22, the CNN inference unit 32 infers the image features identified by the image data acquired from the inference-side input / output unit 31 and obtains feature data as an inference result. The CNN inference unit 32 outputs the obtained feature data to the anomaly distribution calculation unit 33. The anomaly distribution calculation unit 33 acquires the feature data from the CNN inference unit 32 and acquires the learning result (data showing the low-rank matrix) from the low-rank approximation unit 23 of the learning system 20. The anomaly distribution calculation unit 33 calculates the distance between the inferred feature data and the multidimensional space based on the low-rank matrix, based on the feature data obtained from the CNN inference unit 32 and the learning results obtained from the low-rank approximation unit 23, and performs anomaly detection based on the calculated result. The anomaly distribution calculation unit 33 outputs the result of the anomaly detection to, for example, the output interface 5.

[0023] As described above, the low-rank approximation unit 23 calculates the singular value decomposition of features that are identified by the feature data and expressed in matrix form. Here, the low-rank approximation unit 23 calculates the singular value decomposition using the one-sided Jacobi method. The algorithm of the one-sided Jacobi method is to (1) calculate the Gram matrix from a matrix representing the features identified by the feature data (hereinafter also simply referred to as "matrix"), (2) determine convergence based on whether the calculated Gram matrix is ​​approaching a diagonal matrix, (3) calculate the Jacobi rotation matrix, which is a matrix for reducing the off-diagonal components of the Gram matrix, and (4) apply the Jacobi rotation matrix to the matrix and update the matrix. This process is repeated to bring the matrix closer to the singular value decomposed matrix. In (4), applying the Jacobi rotation matrix to the matrix is ​​synonymous with multiplying the matrix by the Jacobi rotation matrix.

[0024] Regarding the calculation of the Gram matrix in (1), the one-sided Jacobi method allows for the efficient determination of the singular values ​​and singular vectors of a matrix by calculating the Gram matrix. The Gram matrix is ​​a symmetric matrix whose elements are the inner products of the column vectors of the matrix, and by using this Gram matrix, the calculation of the singular value decomposition of the matrix is ​​simplified.

[0025] Regarding the calculation of the Jacobian rotation matrix in (3), the Jacobian rotation matrix is ​​used to bring specific elements (specifically, off-diagonal elements) of the Gram matrix closer to zero in order to perform singular value decomposition of the matrix. Specifically, by applying the Jacobian rotation matrix to the matrix, two specific columns (or rows) of the matrix are rotated, and the matrix is ​​updated. This application of the Jacobian rotation matrix and the updating of the matrix are repeated, aiming to bring the off-diagonal elements of the Gram matrix calculated from the updated matrix to zero. As a result, the Gram matrix is ​​diagonalized, and the process of finding the singular values ​​and singular vectors in the singular value decomposition of the matrix proceeds.

[0026] The Jacobian rotation matrix is ​​applied iteratively to the matrix, and the Gram matrix is ​​iteratively calculated from the matrix to which the Jacobian rotation matrix has been applied. This causes the Gram matrix to eventually approach a diagonal matrix. At this time, the calculation of the Gram matrix in (1) and the updating of the matrix in (4) mainly involve matrix calculations (matrix operations), and the computational load is relatively high. Therefore, the data processing device 1 performs these two matrix operations in the matrix operation unit 51 of the NPU 50 until the first convergence criterion is met. Note that the convergence criterion in (2) based on the first convergence criterion is performed in the low-rank approximation unit 23, but the calculation of the Jacobian rotation matrix in (3) until the first convergence criterion is met may be performed by the low-rank approximation unit 23 or by the matrix operation unit 51 of the NPU 50.

[0027] Then, after the first convergence criterion is met, the data processing device 1 performs calculations (1) to (4) in the low-rank approximation unit 23 until it meets a second convergence criterion, which has stricter convergence criteria than the first. In other words, the data processing device 1 performs the singular value decomposition calculation of the matrix representing the features identified by the feature data, with the low-rank approximation unit 23 in the learning system 20 built on the CPU 10 and the matrix operation unit 51 whose function is realized by the NPU 50 working together.

[0028] Next, the detailed functions of the low-rank approximation unit 23 and the matrix calculation unit 51 will be described with reference to an example of their configuration. Figure 3 is a block diagram showing an example of the configuration of the low-rank approximation unit 23 and the matrix calculation unit 51.

[0029] The low-rank approximation unit 23 calculates singular value decomposition of a matrix representing features identified primarily from feature data obtained from the CNN learning unit 22, and compresses the dimensionality of the features by performing low-rank approximation of the feature data based on the results. The low-rank approximation unit 23 also generates a low-rank matrix from the dimensionally compressed features.

[0030] The low-rank approximation unit 23 includes a feature acquisition unit (matrix acquisition unit) 231, an NPU convergence determination unit 232, an NPU Jacobi rotation matrix calculation unit 233, a CPU Gram matrix calculation unit 234, a CPU convergence determination unit 235, a CPU Jacobi rotation matrix calculation unit 236, a CPU matrix update unit 237, and a low-rank processing unit 238. The matrix operation unit 51 includes an NPU Gram matrix calculation unit 511 and an NPU matrix update unit 512. Here, an example is described in which the NPU Jacobi rotation matrix calculation unit 233 is provided in the low-rank approximation unit 23, but the NPU Jacobi rotation matrix calculation unit 233 may also be provided in the matrix operation unit 51.

[0031] The feature acquisition unit 231 acquires feature data from the CNN learning unit 22 and outputs the acquired feature data to the NPU Gram matrix calculation unit 511. At this time, the feature acquisition unit 231 divides the acquired feature data into appropriate sizes for calculation by the NPU Gram matrix calculation unit 511 and outputs the divided feature data to the NPU Gram matrix calculation unit 511.

[0032] To elaborate on the size of the feature data, it has been proposed, for example, to perform singular value decomposition calculations using a GPU. The advantage of using a GPU is that functions for calculating singular value decomposition are already available as libraries. One such library is Nvidia's cuSOLVER library. In this case, the function can process the data quickly if the size of the matrix to be processed is 32 rows x 32 columns or less.

[0033] On the other hand, the neural network processor (NPU 50) can use a matrix with a size larger than 32 rows x 32 columns (for example, 128 rows x 128 columns). Therefore, by using a matrix with a size of 128 rows x 128 columns, the neural network processor does not need to finely divide the input data, i.e., the matrix to be processed, for calculation, and can perform calculations at high speed. Accordingly, the feature acquisition unit 231 divides the feature data acquired from the CNN learning unit 22 into sizes that the NPU Gram matrix calculation unit 511 can calculate (for example, 128 rows x 128 columns), and outputs the divided feature data to the NPU Gram matrix calculation unit 511. As a result, the feature acquisition unit 231 does not need to finely divide the matrix to be processed for output, and can speed up processing. Alternatively, the feature acquisition unit 231 may adjust the size of the feature data acquired from the CNN learning unit 22 to be close to the upper limit of the matrix size that the NPU Gram matrix calculation unit 511 can calculate, and then output the adjusted feature data to the NPU Gram matrix calculation unit 511.

[0034] The NPU Gram matrix calculation unit 511, upon acquiring feature data divided into the above-mentioned sizes from the feature acquisition unit 231, calculates a Gram matrix based on the acquired feature data. The NPU Gram matrix calculation unit 511 outputs the feature data acquired from the feature acquisition unit 231 and data showing the calculated Gram matrix to the NPU convergence determination unit 232. For the sake of simplicity, the data showing the Gram matrix will also be simply referred to as the "Gram matrix" below.

[0035] The NPU convergence determination unit 232 obtains feature data and a Gram matrix from the NPU Gram matrix calculation unit 511. Based on the obtained Gram matrix, the NPU convergence determination unit 232 performs a first convergence determination based on a first convergence determination condition. The first convergence determination condition is, for example, the Gram matrix S as shown in equation (1) below. ij The norm H of the off-diagonal elements (i ≠ j) is half-precision (10 -3 ) until it is less than or equal to the specified value.

[0036] If the NPU convergence determination unit 232 determines, as a result of performing the first convergence determination, that the first convergence determination condition is not met, it outputs the feature data and Gram matrix obtained from the NPU Gram matrix calculation unit 511 to the NPU Jacobian rotation matrix calculation unit 233.

[0037] The NPU Jacobi rotation matrix calculation unit 233 obtains feature data and a Gram matrix from the NPU convergence determination unit 232. Based on the obtained Gram matrix, the NPU Jacobi rotation matrix calculation unit 233 calculates a Jacobi rotation matrix to bring the off-diagonal elements of the Gram matrix closer to zero. Known methods may be used for this calculation. The NPU Jacobi rotation matrix calculation unit 233 outputs the feature data obtained from the NPU convergence determination unit 232 and data showing the calculated Jacobi rotation matrix to the NPU matrix update unit 512. For the sake of simplicity, in the following explanation, the data showing the Jacobi rotation matrix will also be simply referred to as the "Jacobi rotation matrix".

[0038] The matrix updating unit 512 for NPU acquires the feature amount data and the Jacobi rotation matrix from the Jacobi rotation matrix calculating unit 233 for NPU. The matrix updating unit 512 for NPU applies the acquired Jacobi rotation matrix to a matrix indicating feature amounts specified by the acquired feature amount data, and updates the matrix. The matrix updating unit 512 for NPU outputs feature amount data indicating the feature amount represented by the updated matrix to the Gram matrix calculating unit 511 for NPU.

[0039] The Gram matrix calculating unit 511 for NPU acquires feature amount data indicating the feature amount represented by the updated matrix from the matrix updating unit 512 for NPU. The Gram matrix calculating unit 511 for NPU calculates a Gram matrix from the feature amount data acquired from the matrix updating unit 512 for NPU. The Gram matrix calculating unit 511 for NPU outputs the feature amount data acquired from the feature amount acquiring unit 231 and the calculated Gram matrix to the convergence determining unit 232 for NPU.

[0040] The convergence determining unit 232 for NPU acquires the feature amount data and the Gram matrix from the Gram matrix calculating unit 511 for NPU. The convergence determining unit 232 for NPU performs first convergence determination based on the first convergence determination condition described above on the basis of the acquired Gram matrix. When the convergence determining unit 232 for NPU determines that the first convergence determination condition is not satisfied as a result of performing the first convergence determination, it outputs the feature amount data and the Gram matrix acquired from the Gram matrix calculating unit 511 for NPU to the Jacobi rotation matrix calculating unit 233 for NPU in the same manner as described above.

[0041] On the other hand, when the convergence determining unit 232 for NPU determines that the first convergence determination condition is satisfied as a result of performing the first convergence determination, it outputs the feature amount data acquired from the Gram matrix calculating unit 511 for NPU to the Gram matrix calculating unit 234 for CPU. Thereafter, the calculation of singular value decomposition is performed by the low-rank approximation unit 23.

[0042] The Gram matrix calculating unit 234 for CPU acquires feature amount data from the convergence determining unit 232 for NPU. The Gram matrix calculating unit 234 for CPU calculates a Gram matrix from the acquired feature amount data. The Gram matrix calculating unit 234 for CPU outputs the feature amount data acquired from the convergence determining unit 232 for NPU and the calculated Gram matrix to the convergence determining unit 235 for CPU.

[0043] The CPU convergence determination unit 235 obtains feature data and a Gram matrix from the CPU Gram matrix calculation unit 234. Based on the obtained Gram matrix, the CPU convergence determination unit 235 performs a second convergence determination based on a second convergence determination condition. The second convergence determination condition is, for example, the Gram matrix S as shown in equation (2) below. ij The norm H of the off-diagonal elements (i ≠ j) is single precision (10 -7 ) until it is less than or equal to the specified value.

[0044] If the CPU convergence determination unit 235 determines, as a result of performing the second convergence determination, that the second convergence determination condition is not met, it outputs the feature data and Gram matrix obtained from the CPU Gram matrix calculation unit 234 to the CPU Jacobian rotation matrix calculation unit 236.

[0045] The CPU Jacobi rotation matrix calculation unit 236 obtains feature data and a Gram matrix from the CPU convergence determination unit 235. Based on the obtained Gram matrix, the CPU Jacobi rotation matrix calculation unit 236 calculates a Jacobi rotation matrix to bring the off-diagonal elements of the Gram matrix closer to zero. Known methods may be used for this calculation. The CPU Jacobi rotation matrix calculation unit 236 outputs the feature data obtained from the CPU convergence determination unit 235 and the calculated Jacobi rotation matrix to the CPU matrix update unit 237.

[0046] The CPU matrix update unit 237 obtains feature data and a Jacobian rotation matrix from the CPU Jacobian rotation matrix calculation unit 236. The CPU matrix update unit 237 applies the obtained Jacobian rotation matrix to a matrix representing the features identified by the obtained feature data and updates the matrix. The CPU matrix update unit 237 outputs feature data representing the features represented by the updated matrix to the CPU Gram matrix calculation unit 234.

[0047] The CPU Gram matrix calculation unit 234 acquires feature data from the CPU matrix update unit 237. The CPU Gram matrix calculation unit 234 calculates a Gram matrix from the feature data acquired from the CPU matrix update unit 237. The CPU Gram matrix calculation unit 234 outputs the feature data acquired from the CPU matrix update unit 237 and the calculated Gram matrix to the CPU convergence determination unit 235.

[0048] The CPU convergence determination unit 235 obtains feature data and the Gram matrix from the CPU Gram matrix calculation unit 234. Based on the obtained Gram matrix, the CPU convergence determination unit 235 performs a second convergence determination based on the second convergence determination condition described above. If the NPU convergence determination unit 232 determines that the second convergence determination condition is not met as a result of the second convergence determination, it outputs the feature data and the Gram matrix obtained from the CPU Gram matrix calculation unit 234 to the CPU Jacobian rotation matrix calculation unit 236, as described above. On the other hand, if the CPU convergence determination unit 235 determines that the second convergence determination condition is met as a result of the second convergence determination, it terminates the singular value decomposition calculation and outputs the feature data obtained from the CPU Gram matrix calculation unit 234 to the low-rank processing unit 238.

[0049] The low-rank processing unit 238 acquires feature data from the CPU convergence determination unit 235. At this point, the result of singular value decomposition of a matrix representing the features identified by the feature data is obtained, and the low-rank processing unit 238 compresses the dimensionality of the features by performing a low-rank approximation of the feature data based on the result of the singular value decomposition. The low-rank processing unit 238 also generates a low-rank matrix from the dimensionally compressed features and outputs data representing the generated low-rank matrix as a learning result to the anomaly distribution calculation unit 33 of the inference system 30.

[0050] Next, an example of the operation of the low-rank approximation unit 23 and the matrix calculation unit 51 in the data processing device 1 according to Embodiment 1 will be described. Figure 4 is a flowchart showing an example of the operation of the low-rank approximation unit 23 and the matrix calculation unit 51 in the data processing device 1.

[0051] First, the feature acquisition unit 231 acquires feature data from the CNN learning unit 22 (step ST1). The feature acquisition unit 231 divides the acquired feature data into appropriate sizes for calculation by the NPU Gram matrix calculation unit 511, and outputs the divided feature data to the NPU Gram matrix calculation unit 511.

[0052] When the NPU Gram matrix calculation unit 511 obtains feature data divided into the above-mentioned sizes from the feature acquisition unit 231, it calculates a Gram matrix based on the acquired feature data (step ST2). The NPU Gram matrix calculation unit 511 outputs the feature data obtained from the feature acquisition unit 231 and the calculated Gram matrix to the NPU convergence determination unit 232.

[0053] The NPU convergence determination unit 232 obtains feature data and a Gram matrix from the NPU Gram matrix calculation unit 511. Based on the obtained Gram matrix, the NPU convergence determination unit 232 performs a first convergence determination based on a first convergence determination condition (step ST3). The first convergence determination condition is, for example, the Gram matrix S as shown in equation (1) above. ij The norm H of the off-diagonal elements (i ≠ j) is half-precision (10 -3 ) until it is less than or equal to the specified value.

[0054] If the NPU convergence determination unit 232 determines, as a result of the first convergence determination, that the first convergence determination condition is not met (step ST3; No), it outputs the feature data and Gram matrix obtained from the NPU Gram matrix calculation unit 511 to the NPU Jacobian rotation matrix calculation unit 233. The process then proceeds to step ST4. On the other hand, if the NPU convergence determination unit 232 determines, as a result of the first convergence determination, that the first convergence determination condition is met (step ST3; Yes), it outputs the feature data obtained from the NPU Gram matrix calculation unit 511 to the CPU Gram matrix calculation unit 234. The process then proceeds to step ST6.

[0055] In step ST4, the NPU Jacobi rotation matrix calculation unit 233 obtains feature data and a Gram matrix from the NPU convergence determination unit 232. Based on the obtained Gram matrix, the NPU Jacobi rotation matrix calculation unit 233 calculates a Jacobi rotation matrix to bring the off-diagonal components of the Gram matrix closer to zero (step ST4). The NPU Jacobi rotation matrix calculation unit 233 outputs the feature data obtained from the NPU convergence determination unit 232 and the calculated Jacobi rotation matrix to the NPU matrix update unit 512.

[0056] The NPU matrix update unit 512 obtains feature data and a Jacobian rotation matrix from the NPU Jacobian rotation matrix calculation unit 233. The NPU matrix update unit 512 applies the obtained Jacobian rotation matrix to a matrix representing the features identified by the obtained feature data and updates the matrix (step ST5). The NPU matrix update unit 512 outputs feature data representing the features represented by the updated matrix to the NPU Gram matrix calculation unit 511. After that, the process returns to step ST2.

[0057] In step ST6, the CPU Gram matrix calculation unit 234 acquires feature data from the NPU convergence determination unit 232. The CPU Gram matrix calculation unit 234 calculates a Gram matrix based on the acquired feature data (step ST6). The CPU Gram matrix calculation unit 234 outputs the feature data acquired from the NPU convergence determination unit 232 and the calculated Gram matrix to the CPU convergence determination unit 235.

[0058] The CPU convergence determination unit 235 obtains feature data and a Gram matrix from the CPU Gram matrix calculation unit 234. Based on the obtained Gram matrix, the CPU convergence determination unit 235 performs a second convergence determination based on a second convergence determination condition (step ST7). The second convergence determination condition is, for example, the Gram matrix S as shown in equation (2) above. ij The norm H of the off-diagonal elements (i ≠ j) is single precision (10 -7 ) until it is less than or equal to the specified value.

[0059] If the CPU convergence determination unit 235 determines, as a result of the second convergence determination, that the second convergence determination condition is not met (step ST7; No), it outputs the feature data and Gram matrix obtained from the CPU Gram matrix calculation unit 234 to the CPU Jacobian rotation matrix calculation unit 236. The process then proceeds to step ST8. On the other hand, if the CPU convergence determination unit 235 determines, as a result of the second convergence determination, that the second convergence determination condition is met (step ST7; Yes), it terminates the singular value decomposition calculation and outputs the feature data obtained from the CPU Gram matrix calculation unit 234 to the low-rank processing unit 238. The process then proceeds to step ST10.

[0060] In step ST8, the CPU Jacobi rotation matrix calculation unit 236 obtains feature data and a Gram matrix from the CPU convergence determination unit 235. Based on the obtained Gram matrix, the CPU Jacobi rotation matrix calculation unit 236 calculates a Jacobi rotation matrix to bring the off-diagonal components of the Gram matrix closer to zero (step ST8). The CPU Jacobi rotation matrix calculation unit 236 outputs the feature data obtained from the CPU convergence determination unit 235 and the calculated Jacobi rotation matrix to the CPU matrix update unit 237.

[0061] The CPU matrix update unit 237 obtains feature data and the Jacobian rotation matrix from the CPU Jacobian rotation matrix calculation unit 236. The CPU matrix update unit 237 applies the obtained Jacobian rotation matrix to a matrix representing the features identified by the obtained feature data and updates the matrix (step ST9). The CPU matrix update unit 237 outputs feature data representing the features represented by the updated matrix to the CPU Gram matrix calculation unit 234. After that, the process returns to step ST6.

[0062] In step ST10, the rank reduction processing unit 238 acquires feature amount data from the CPU convergence determination unit 235. At this point, a result of singular value decomposition of a matrix indicating the feature amount specified by the feature amount data is obtained, so the rank reduction processing unit 238 compresses the dimension of the feature amount by performing low-rank approximation of the feature amount data from the result of the singular value decomposition. Further, the rank reduction processing unit 238 generates a low-rank matrix from the dimensionally compressed feature amount, and outputs data indicating the generated low-rank matrix to the abnormal distribution calculation unit 33 of the inference system 30 as a learning result (step ST10).

[0063] Note that the first convergence determination condition and the second convergence determination condition described above are merely examples, and conditions other than those described above may be used. For example, the first convergence determination condition is that the Gram matrix S ij the maximum value of the off-diagonal elements (i≠j) is single-precision (10 -7 ) It may be set until the value becomes equal to or less than the above. Further, in this case, the second convergence determination condition is that the Gram matrix S ij the maximum value of the off-diagonal elements (i≠j) is double-precision (10 -15 ) It may be set until the value becomes equal to or less than the above.

[0064] As described above, the data processing device 1 calculates singular value decomposition of a matrix indicating a feature amount specified by feature amount data, and performs the calculation of this singular value decomposition by the one-sided Jacobi method. The algorithm of the one-sided Jacobi method is as follows: (1) calculate a Gram matrix from a matrix (original matrix) indicating the feature amount specified by the feature amount data; (2) perform convergence determination based on whether the calculated Gram matrix is approaching a diagonal matrix; (3) calculate a Jacobi rotation matrix, which is a matrix for reducing off-diagonal components of the Gram matrix; (4) apply the Jacobi rotation matrix to the original matrix and update the original matrix. This processing is repeated to bring the original matrix closer to a singular value decomposed matrix.

[0065] At this time, the calculation of the Gram matrix in (1) and the updating of the matrix in (4) mainly concern matrix calculations (matrix operations), and the data processing device 1 iteratively performs these two matrix operations using the neural network processor (specifically, the matrix operation unit 51) until the first convergence criterion is met. After the first convergence criterion is met, the data processing device 1 iteratively performs the calculations (1) to (4), including (1) and (4) above, using the general-purpose processor (specifically, the low-rank approximation unit 23) until the second convergence criterion, which has stricter convergence criterion conditions than the first, is met. At this time, the choice of which processor performs the matrix operations switches when the first convergence criterion is met. From the start of the singular value decomposition calculation until the first convergence criterion is met, the neural network processor, which is good at matrix operations, is the main processor for matrix operations. After the first convergence criterion is met, until the second convergence criterion is met, the general-purpose processor is the main processor for matrix operations, thereby improving the accuracy of the matrix operations.

[0066] In other words, the accuracy of matrix operations performed by a neural network processor is expected to be lower than that of a general-purpose processor. However, the low-precision calculations performed by the neural network processor are solely for providing initial estimations, and their purpose is to improve the initial values ​​that serve as input for the high-precision calculations performed by the general-purpose processor. Even if the initial estimations performed by the neural network processor are somewhat inaccurate, these errors will be corrected during the iterative process of high-precision calculations performed by the general-purpose processor. Furthermore, many iterative methods guarantee convergence if the initial estimations are accurate to a certain extent. That is, if the initial estimations obtained from the low-precision calculations performed by the neural network processor are within a certain range of accuracy, the calculation results will converge through the final high-precision calculations performed by the general-purpose processor.

[0067] Furthermore, in a neural network processor, the size of the matrix to be processed can be set to a size larger than the upper limit of the matrix size processed by the function used for calculating singular value decomposition, which is available on the GPU (for example, 32 rows x 32 columns) (for example, 128 rows x 128 columns). Therefore, in the data processing device 1, by setting the size of the matrix to be processed by the neural network processor to, for example, 128 rows x 128 columns, it is not necessary to divide the input data, i.e., the size of the matrix to be processed, into smaller parts for calculation, and calculations can be performed at high speed. Alternatively, the data processing device 1 may adjust the size of the matrix to be processed by the neural network processor to be close to the upper limit of the matrix size that the neural network processor can calculate, and even in this case, calculations can be performed at high speed.

[0068] As described above, the data processing device 1 according to Embodiment 1 performs singular value decomposition calculations in cooperation with a small, low-power neural network processor and a general-purpose processor capable of high-precision calculations. This enables both power saving and miniaturization of the device, as well as performance and accuracy of the singular value decomposition calculation process. Furthermore, the data processing device 1 according to Embodiment 1 can increase the size of the data input to the neural network processor, enabling both increased capacity of feature data input by the neural network processor and improved calculation speed of singular value decomposition.

[0069] (Example of application) The data processing system including the data processing device 1 according to Embodiment 1 can be applied, for example, to the visual inspection described above. Visual inspection is an inspection that checks the appearance of a part or product in order to maintain or guarantee the quality of the part or product, and mainly checks for visual defects such as foreign matter, dirt, scratches, burrs, chips, or deformation attached to the surface of the part or product, and makes an abnormality determination of the part or product based on the result. For example, the sensor 60 acquires data showing an image of the appearance of the part or product to be inspected (hereinafter referred to as "part to be inspected") as the image data to be acquired. At this time, the sensor 60 acquires data showing a normal image, which is an image of the part to be inspected in a state without abnormalities, as the image to be used in the learning system 20. The learning-side input / output unit 21 of the learning system 20 acquires data showing a normal image from the sensor 60 and outputs the acquired data to the CNN learning unit 22. Subsequently, the learning system 20 uses the above method to compress the feature quantities of the normal image to generate a low-rank matrix, and outputs data showing the generated low-rank matrix as a learning result to the abnormality distribution calculation unit 33 of the inference system 30. Meanwhile, the sensor 60 acquires data representing an inspection image, which is an image of the part to be inspected during the actual inspection. The inference-side input / output unit 31 of the inference system 30 acquires data representing the inspection image from the sensor 60 and outputs the acquired data to the CNN inference unit 32. Subsequently, the inference system 30 uses the above method to infer the features of the inspection image, obtains feature data as an inference result, calculates the distance between the feature data and the multidimensional space based on the low-rank matrix, and makes an abnormality determination for the part to be inspected based on the calculated result. In this case, the data processing system including the data processing device 1 according to Embodiment 1 functions as an inspection system that performs visual inspection of the part to be inspected. Even in this case, the data processing device 1 can achieve both power saving and miniaturization of the device and performance and accuracy of the singular value decomposition calculation process by having a small, low-power neural network processor and a general-purpose processor capable of high-precision calculations work together to perform singular value decomposition calculations.Furthermore, the data processing device 1 can increase the size of the data input to the neural network processor, enabling both increased capacity for feature data input by the neural network processor and improved calculation speed for singular value decomposition. In addition, the inspection system including the data processing device 1 can perform abnormality detection of the inspected part at high speed and with high accuracy. Moreover, the data processing system including the data processing device 1 according to Embodiment 1 is not limited to the visual inspection described above, but can be broadly applied to situations such as performing abnormality detection of the inspected part based on the distance (Mahalanobis distance) between the feature quantities obtained from inspection images taken of the inspected part and the feature quantity distribution of normal images learned by a convolutional neural network (CNN), etc.

[0070] As described above, according to this embodiment 1, the data processing device comprises a general-purpose processor and a neural network processor, and is a data processing device 1 that performs singular value decomposition calculation on a predetermined matrix using the one-sided Jacobi method. From the start of the singular value decomposition calculation until the first convergence criterion is met, the neural network processor performs predetermined matrix operations that are iteratively performed in the calculation process of singular value decomposition using the one-sided Jacobi method. From the time the first convergence criterion is met until the second convergence criterion, which has stricter convergence criterion conditions than the first convergence criterion, is met, the general-purpose processor performs predetermined matrix operations. As a result, the data processing device 1 according to embodiment 1 can achieve both power saving and miniaturization of the device that performs the singular value decomposition calculation, and the performance and accuracy of the calculation process.

[0071] Furthermore, the predetermined matrix operation includes calculating a Gram matrix based on the matrix, and updating the matrix by applying the Jacobian rotation matrix calculated based on the calculated Gram matrix to the matrix. As a result, the data processing device 1 according to Embodiment 1 can perform matrix operations, which are particularly computationally intensive, through cooperation between a general-purpose processor and a neural network processor.

[0072] Furthermore, when the neural network processor starts the singular value decomposition calculation, it receives input data representing a predetermined matrix, the matrix size being adjusted to approach the upper limit of the matrix size that the neural network processor can process. As a result, the data processing device 1 according to Embodiment 1 can increase the amount of data input to the neural network processor and speed up processing.

[0073] Furthermore, the general-purpose processor is composed of a CPU 10, and the neural network processor is composed of an NPU 50. The NPU 50 implements the functions of an NPU Gram matrix calculation unit 511 that acquires data representing a matrix and calculates a Gram matrix from the matrix identified by the acquired data, and an NPU matrix update unit 512 that updates the Gram matrix by applying a Jacobi rotation matrix to the Gram matrix calculated by the NPU Gram matrix calculation unit 511 to bring the off-diagonal components of the Gram matrix closer to zero, and outputs data representing the updated matrix to the NPU Gram matrix calculation unit 511. The CPU 10 acquires data representing a predetermined matrix prior to the start of singular value decomposition calculations. The system implements the functions of the following units: a matrix acquisition unit 231 that acquires data, adjusts the size of the matrix indicated by the acquired data so that it approaches the upper limit of the matrix size that the NPU 50 can process, and then outputs it to the NPU 50; a CPU Gram matrix calculation unit 234 that acquires data indicating a matrix and calculates a Gram matrix from the matrix identified by the acquired data; and a CPU matrix update unit 237 that updates the Gram matrix by applying a Jacobi rotation matrix to the Gram matrix calculated by the CPU Gram matrix calculation unit 234 to bring the off-diagonal components of the Gram matrix closer to zero, and then outputs data indicating the updated matrix to the CPU Gram matrix calculation unit 234.

[0074] Furthermore, the CPU 10 implements the functions of the following units: an NPU convergence determination unit 232 that acquires the Gram matrix calculated by the NPU Gram matrix calculation unit 511 and performs a first convergence determination based on a first convergence determination condition based on the acquired Gram matrix; an NPU Jacobi rotation matrix calculation unit 233 that, if the NPU convergence determination unit 232 determines that the first convergence determination condition is not met, calculates a Jacobi rotation matrix to bring the off-diagonal components of the Gram matrix calculated by the NPU Gram matrix calculation unit 511 closer to zero; a CPU convergence determination unit 235 that acquires the Gram matrix calculated by the CPU Gram matrix calculation unit 234 and performs a second convergence determination based on a second convergence determination condition based on the acquired Gram matrix; and a CPU Jacobi rotation matrix calculation unit 236 that, if the CPU convergence determination unit 235 determines that the second convergence determination condition is not met, calculates a Jacobi rotation matrix to bring the off-diagonal components of the Gram matrix calculated by the CPU Gram matrix calculation unit 234 closer to zero. As a result, the data processing device 1 according to Embodiment 1 can appropriately perform the first and second convergence determinations and the calculation of the Jacobian rotation matrix.

[0075] Although preferred embodiments have been described in detail above, the invention is not limited to the embodiments described above, and various modifications and substitutions can be made to the embodiments described above without departing from the scope of the claims.

[0076] Furthermore, this disclosure allows for modifications of any component of the embodiments, or the omission of any component in each embodiment.

[0077] This disclosure makes it possible to achieve both power saving and miniaturization of a device that performs singular value decomposition calculations, and the performance and accuracy of said calculations, and is suitable for use in data processing devices and data processing methods.

[0078] 1 Data processing unit, 2 Processor, 3 Memory, 4 Input interface, 5 Output interface, 10 CPU, 20 Learning system, 21 Learning-side input / output unit, 22 CNN learning unit, 23 Low-rank approximation unit, 30 Inference system, 31 Inference-side input / output unit, 32 CNN inference unit, 33 Anomaly distribution calculation unit, 51 Matrix calculation unit, 60 Sensor, 231 Feature acquisition unit (matrix acquisition unit), 232 NPU convergence determination unit, 233 NPU Jacobi rotation matrix calculation unit, 234 CPU Gram matrix calculation unit, 235 CPU convergence determination unit, 236 CPU Jacobi rotation matrix calculation unit, 237 CPU matrix update unit, 238 Low-rank processing unit, 511 NPU Gram matrix calculation unit, 512 NPU matrix update unit.

Claims

1. A data processing device comprising a general-purpose processor and a neural network processor, which performs singular value decomposition calculation on a predetermined matrix using the one-sided Jacobi method, wherein the neural network processor performs predetermined matrix operations that are iteratively performed in the calculation process of singular value decomposition using the one-sided Jacobi method from the start of the singular value decomposition calculation until a first convergence criterion is met, and the general-purpose processor performs the predetermined matrix operations from the time the first convergence criterion is met until a second convergence criterion, which has stricter convergence criterion conditions than the first convergence criterion, is met.

2. The data processing apparatus according to claim 1, characterized in that the predetermined matrix operation includes calculating a Gram matrix based on the matrix, and updating the matrix by applying a Jacobi rotation matrix calculated based on the calculated Gram matrix to the matrix.

3. The data processing device according to claim 1 or 2, characterized in that, when the calculation of the singular value decomposition of the neural network is started, the neural network processor receives input of data representing the predetermined matrix, wherein the size of the matrix is ​​adjusted to approach the upper limit of the size of a matrix that the neural network processor can process.

4. The general-purpose processor is composed of a CPU, the neural network processor is composed of an NPU, the NPU implements the functions of an NPU Gram matrix calculation unit which acquires data representing the matrix and calculates a Gram matrix from the matrix specified by the acquired data, and an NPU matrix update unit which updates the Gram matrix by applying a Jacobi rotation matrix to the Gram matrix calculated by the NPU Gram matrix calculation unit to bring the off-diagonal components of the Gram matrix to near zero, and outputs data representing the updated matrix to the NPU Gram matrix calculation unit, the CPU implements the functions of each unit, the CPU implements the functions of an NPU Gram matrix calculation unit which acquires data representing the predetermined matrix prior to the start of the singular value decomposition calculation, adjusts the size of the matrix represented by the acquired data so that the size of the matrix represented by the acquired data approaches the upper limit of the size of a matrix that the NPU can process, and outputs it to the NPU, and a CPU Gram matrix calculation unit which acquires data representing the matrix and calculates a Gram matrix from the matrix specified by the acquired data, A data processing device according to any one of claims 1 to 3, characterized in that it realizes the functions of each part, including a CPU matrix update unit that updates the Gram matrix calculated by the CPU Gram matrix calculation unit by applying a Jacobi rotation matrix to the Gram matrix to bring the off-diagonal elements of the Gram matrix closer to zero, and outputs data showing the updated matrix to the CPU Gram matrix calculation unit.

5. The data processing device according to claim 4, characterized in that the CPU implements the functions of each of the following units: an NPU convergence determination unit that acquires a Gram matrix calculated by the NPU Gram matrix calculation unit and performs a first convergence determination based on the first convergence determination condition based on the acquired Gram matrix; an NPU Jacobi rotation matrix calculation unit that calculates a Jacobi rotation matrix to bring the off-diagonal components of the Gram matrix calculated by the NPU Gram matrix calculation unit closer to zero if the NPU convergence determination unit determines that the first convergence determination condition is not met; a CPU convergence determination unit that acquires a Gram matrix calculated by the CPU Gram matrix calculation unit and performs a second convergence determination based on the second convergence determination condition based on the acquired Gram matrix; and a CPU Jacobi rotation matrix calculation unit that calculates a Jacobi rotation matrix to bring the off-diagonal components of the Gram matrix calculated by the CPU Gram matrix calculation unit closer to zero if the CPU convergence determination unit determines that the second convergence determination condition is not met.

6. A data processing method comprising a data processing device that performs singular value decomposition calculation on a predetermined matrix using the one-sided Jacobi method, wherein the neural network processor performs predetermined matrix operations that are iteratively performed in the process of calculating singular value decomposition using the one-sided Jacobi method from the start of the singular value decomposition calculation until a first convergence criterion is met, and the general-purpose processor performs the predetermined matrix operations from the time the first convergence criterion is met until a second convergence criterion, which has stricter convergence criterion conditions than the first convergence criterion, is met.