Information Processing Program, Information Processing Method, and Information Processing Apparatus

By grouping non-zero elements in multi-dimensional tensor data to facilitate parallel processing, the method addresses the inefficiencies in MTTKRP operations, reducing calculation time and memory usage while ensuring accurate results.

JP7698197B2Active Publication Date: 2025-06-25FUJITSU LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2021136692
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-08-24
Publication Date
2025-06-25
Estimated Expiration
2041-08-24

AI Technical Summary

Technical Problem

Existing methods for analyzing multi-dimensional data, such as ALS, face challenges with increased calculation time and memory usage due to handling large matrices in MTTKRP operations, and parallel processing often results in conflicting operations.

Method used

An information processing method that identifies combinations of non-zero elements in multi-dimensional tensor data, groups them to avoid index overlaps, and performs parallel processing within these groups to reduce calculation time and memory usage.

Benefits of technology

The method significantly reduces calculation time and memory requirements for MTTKRP processing, eliminating bottlenecks in ALS and ensuring accurate results by avoiding conflicts between operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698197000001
    Figure 0007698197000001
  • Figure 0007698197000002
    Figure 0007698197000002
  • Figure 0007698197000003
    Figure 0007698197000003
Patent Text Reader

Abstract

To reduce calculation time.SOLUTION: An information processing apparatus 100 acquires first data 101 that enables, for each of non-zero elements included in multidimensional tensor data, specification of a combination of a value of the element and an index of each dimension that indicates a position of the element. The information processing apparatus 100 generates, on the basis of the acquired first data 101, second data 102 that enables specification of a plurality of groups obtained by grouping each of the combinations such that the combinations with indexes that overlap with each other are included in different groups. The information processing apparatus 100 performs, on the basis of the generated second data 102, MTTKRP processing by setting each combination of a plurality of combinations included in the group as a target of parallel processing in the MTTKRP processing related to the tensor data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing program, an information processing method, and an information processing apparatus.

Background Art

[0002] Conventionally, there is a technique for analyzing the characteristics of multidimensional data by decomposing the multidimensional data called a tensor. For example, there is a technique called ALS (Alternating Least Squares) for analyzing the characteristics of multidimensional data.

[0003] As prior art, for example, there is one that performs loop calculations for a plurality of indexes of tensor data on N-dimensional tensor data. Also, for example, there is a technique for updating a plurality of factor matrices as a set of factor matrices capable of parallel processing. Also, for example, there is a technique for creating a partitioned compressed representation of a matrix including a mapping array including a list. Also, for example, there is a technique for extracting a pair of the value of a non-zero element and a column identifier from each first row included in a matrix, and adding a dummy pair having a value of zero to the first row in which the number of non-zero elements is less than the maximum value.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the prior art, there is a problem that the calculation time increases when decomposing multi-dimensional data. For example, in the calculation of ALS, in the calculation called MTTKRP (Matricized Tensor Times Khatri-Rao Product), huge matrices are handled, which tends to increase the calculation time and the memory usage. Also, when trying to reduce the calculation time by parallelizing matrix operations, different operations may conflict with each other.

[0006] In one aspect, the present invention aims to shorten the calculation time for MTTKRP processing.

Means for Solving the Problem

[0007] According to one embodiment, for each non-zero element included in multi-dimensional tensor data, first data that enables identification of a combination of the value of the element and the indexes of each dimension indicating the position of the element is obtained, and based on the obtained first data, second data that enables identification of a plurality of groups obtained by grouping each of the combinations such that the combinations with overlapping indexes are included in different groups is generated, and based on the generated second data, each combination of the plurality of combinations included in the group is used as an object of parallel processing in MTTKRP processing related to the tensor data, and an information processing program, an information processing method, and an information processing apparatus for performing the MTTKRP processing are proposed.

[0008] Also, according to one embodiment, for each non-zero element included in the multi-dimensional tensor data, first data is obtained that enables identification of a combination of the value of the element and the indices of each dimension indicating the position of the element. Based on the obtained first data, the combinations in which the index of the target dimension is discontinuous are not included in the same group according to a predetermined order regarding the index of the target dimension, and second data is generated that enables identification of a plurality of groups for a predetermined number of parallel processes, obtained by grouping each of the combinations. Based on the generated second data, the plurality of groups are used as targets for parallel processing in MTTKRP (Matricized Tensor Times Khatri-Rao Product) processing regarding the tensor data for the target dimension, and operations are performed on each combination included in the group according to the predetermined order, and the results of the operations are stored in a temporary area of the group. Each time an operation regarding one or more combinations in which the index of the target dimension included in the group is the same is completed, the content of the temporary area of the group is reflected in the matrix, thereby implementing an information processing program, an information processing method, and an information processing apparatus for the MTTKRP processing.

Advantages of the Invention

[0009] According to one aspect, it becomes possible to shorten the calculation time for MTTKRP processing.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, with reference to the drawings, embodiments of an information processing program, an information processing method, and an information processing apparatus according to the present invention will be described in detail.

[0012] (An Example of an Information Processing Method According to an Embodiment) FIG. 1 is an explanatory diagram showing an example of an information processing method according to an embodiment. The information processing apparatus 100 is a computer for facilitating the analysis of tensor data. The information processing apparatus 100 is, for example, a server or a PC (Personal Computer).

[0013] Tensor data is multi-dimensional data. Tensor data includes a plurality of elements. The elements are associated with the indices of each dimension. Tensor data may, for example, have sparsity. Sparsity is a property in which among the plurality of elements included in tensor data, the number of zero elements is relatively large and the number of non-zero elements is relatively small.

[0014] Here, there may be cases where it is desired to analyze the characteristics of tensor data. Conventionally, there is a method of analyzing the characteristics of tensor data by decomposing the tensor data. For example, there is a method called CPD (Canonical Polyadic Decomposition).

[0015] Specifically, assume that the tensor data X is of rank 3 and X = R (I × J × K) Let's assume that matrix A = R (I × F) is defined, matrix B = R (J × F) is defined, and matrix C = R (K × F) is defined. Then, X is decomposed into matrix A, matrix B, and matrix C so as to satisfy the following formula (1). In the following description, for convenience, the operation symbol of the Khatri-Rao product may be denoted as "◎".

[0016] X (1) = A(C◎B) T , X (2) = B(C◎A) T , X (3) = C(A◎B) T , ···(1)

[0017] One type of CPD is a method called ALS. ALS is an iterative method. ALS is a method of solving the following formula (2) for updating matrix A by fixing matrix B and matrix C, for example. Specifically, ALS will perform the calculation process shown in formula (3) below.

[0018] A^=min A ||X (1) -A^(C◎B) T || F 2 ···(2)

[0019] A^=X (1) (C◎B)(C T C*B T B) ···(3)

[0020] Among the above formula (3), the calculation of (C T C*B T B) is a dense matrix product. Since the matrix size of (C T C*B T B) is relatively small, the calculation time is considered to be relatively short. On the other hand, among the above formula (3), the calculation of X (1) (C◎B) is such that (C◎B) is R (JK × F) and the matrix size is likely to be relatively large, so it is likely to cause an increase in calculation time and memory usage. The calculation of X (1) (C◎B) is called MTTKRP.

[0021] Similarly, when solving the formula for updating matrix B, the calculation of X (2) (C◎A) will be performed, and the calculation of X (2) (C◎A) is likely to cause an increase in calculation time and memory usage and is called MTTKRP. Similarly, when solving the formula for updating matrix C, the calculation of X (3) (A◎B) will be performed, and the calculation of X (3)The calculation of (A◎B) tends to increase the calculation time and memory usage, and is called MTTKRP. Thus, there is a problem that each MTTKRP tends to be a bottleneck of ALS.

[0022] On the other hand, since X has sparsity, some elements of (C◎B) may not need to be used in the calculation of MTTKRP. Specifically, for the elements of (C◎B) that are multiplied by the zero elements of X, since it is clear that the product of multiplying by the zero elements of X is zero, they may not need to be used in the calculation of MTTKRP.

[0023] Therefore, it is considered that by utilizing sparsity, the calculation of MTTKRP can be carried out based on the non-zero elements of X and the elements of (C◎B) corresponding to the non-zero elements of X without generating the entire (C◎B). Specifically, by the following formula (4), focusing on the f-th column of matrices A, B, and C, the calculation of MTTKRP can be carried out. i, j, k are the indices of each dimension of the non-zero elements of X. val is the value of the non-zero element of X corresponding to the combination of the indices of each dimension of i, j, k.

[0024] A[i][f]+=val*C[k][f]*B[j][f] ···(4)

[0025] Similarly, since X has sparsity, some elements of (C◎A) and (A◎B) may not need to be used in the calculation of MTTKRP. Specifically, for the elements of (C◎A) and (A◎B) that are multiplied by the zero elements of X, since it is clear that the product of multiplying by the zero elements of X is zero, they may not need to be used in the calculation of MTTKRP.

[0026] Furthermore, it may be considered to execute a plurality of operations in the calculation of MTTKRP in parallel to shorten the calculation time. However, in this case, it is difficult to execute a plurality of operations in parallel. For example, different operations may conflict with each other, and the calculation result of MTTKRP may be incorrect. Specifically, X (1)When performing the calculation of (C◎B), for the same A[i][f], different operations may conflict with each other, and the values may not be properly added to A[i][f].

[0027] Also, it is conceivable that X is handled in a format called the CSF (Compressed Sparse Fiber) format in order to reduce the calculation time and memory usage. In this case as well, the problem that different operations conflict with each other and the calculation result of MTTKRP may be incorrect may not be solved.

[0028] Specifically, for any dimension, the problem that different operations conflict with each other and the calculation result of MTTKRP may be incorrect may not be solved. More specifically, X (1) (the calculation of (C◎B) and X (2) (the calculation of (C◎A) and X (3) In the calculation of at least any one of MTTKRP among (the calculation of (A◎B)), it is impossible to avoid different operations from conflicting with each other.

[0029] Thus, conventionally, when analyzing tensor data by CPD, there is a problem of causing an increase in the calculation time and memory usage.

[0030] Therefore, in the present embodiment, an information processing method that can shorten the calculation time and reduce the memory usage when analyzing tensor data by CPD will be described.

[0031] (1-1) The information processing apparatus 100 acquires first data 101. The first data 101 is data that can identify, for each non-zero element included in multi-dimensional tensor data, a combination of the value of the element and the indices of each dimension indicating the position of the element. The first data 101 is, for example, data in COO (COOrdinated) format. The COO format is a format that associates and shows, for each non-zero element, the indices of each dimension of the element and the value of the element. The first data 101 may be, for example, data in CSF format. The information processing apparatus 100 acquires, for example, by acquiring multi-dimensional tensor data and generating the first data 101 based on the acquired multi-dimensional tensor data.

[0032] In the example of FIG. 1, the information processing apparatus 100 acquires first data 101 in COO format corresponding to 3-dimensional tensor data. Each dimension is, for example, I, J, K. Specifically, the first data 101 associates and shows, for each non-zero element included in the 3-dimensional tensor data, the indices i, j, k of each dimension of the element and the value val of the element. The index i corresponds to dimension I. The index j corresponds to dimension J. The index k corresponds to dimension K. Specifically, the information processing apparatus 100 acquires by acquiring 3-dimensional tensor data and generating the first data 101 in COO format based on the acquired 3-dimensional tensor data.

[0033] (1-2) The information processing apparatus 100 generates second data 102 based on the acquired first data 101. The second data 102 is data that can identify a plurality of groups obtained by grouping each combination so that combinations with overlapping indices are included in different groups.

[0034] In the example of FIG. 1, specifically, the indices of the elements of val "a" and the indices of the elements of val "b" overlap. For this reason, as shown by reference numeral 110, the operations on the elements of val "a" and the operations on the elements of val "b" may conflict. Therefore, it is preferable that the combinations related to the elements of val "a" and the combinations related to the elements of val "b" are included in different groups, respectively. Thereby, the information processing apparatus 100 can identify whether the operations on each of any two or more elements can be processed in parallel without conflict.

[0035] (1-3) Based on the generated second data 102, the information processing apparatus 100 performs the MTTKRP process, taking each of the plurality of combinations included in the group as an object of parallel processing in the MTTKRP process for the tensor data. Thereby, in the MTTKRP process, the information processing apparatus 100 can parallel-process different operations while avoiding conflicts between different operations for each dimension. For this reason, the information processing apparatus 100 can reduce the calculation time and memory usage amount related to the MTTKRP process. Also, the information processing apparatus 100 can avoid conflicts between different operations and accurately perform the MTTKRP process.

[0036] Here, the case where the information processing apparatus 100 performs the MTTKRP process based on the second data 102 has been described, but it is not limited thereto. For example, the information processing apparatus 100 may perform the MTTKRP process based on a third data different from the second data 102 instead of the second data 102. The third data is data that can identify a plurality of groups for a predetermined number of parallelisms obtained by grouping each of the combinations so that combinations with discontinuous indices of the target dimension are not included in the same group. An example of the information processing apparatus 100 performing the MTTKRP process based on the third data will be specifically described later with reference to FIGS. 21 to 34.

[0037] Here, the case where the information processing apparatus 100 operates alone has been described, but it is not limited thereto. For example, although the case where the information processing apparatus 100 acquires the first data 101 by generating it in its own apparatus, generates the second data 102, and performs the MTTKRP process has been described, it is not limited thereto.

[0038] For example, the information processing apparatus 100 may cooperate with other computers. Specifically, the information processing apparatus 100 may acquire the first data 101 by receiving it from another computer that generates the first data 101. Specifically, the information processing apparatus 100 may control another computer to perform the MTTKRP process by transmitting the second data 102 to another computer that performs the MTTKRP process. The case where the information processing apparatus 100 cooperates with other computers will be described later with reference to FIG. 2, for example.

[0039] (An example of the CPD calculation system 200) Next, an example of the CPD calculation system 200 to which the information processing apparatus 100 shown in FIG. 1 is applied will be described with reference to FIG. 2.

[0040] FIG. 2 is an explanatory diagram showing an example of the CPD calculation system 200. In FIG. 2, the CPD calculation system 200 includes an information processing apparatus 100, a client apparatus 201, and a calculation apparatus 202.

[0041] In the CPD calculation system 200, the information processing apparatus 100 and the client apparatus 201 are connected via a wired or wireless network 210. The network 210 is, for example, a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, or the like. Also, the information processing apparatus 100 and the calculation apparatus 202 are connected via a wired or wireless network 210.

[0042] The CPD calculation system 200 provides a CPD calculation service to system users. The client device 201 is a computer used by the system user. Based on the operation input of the system user, the client device 201 transmits a request to execute CPD calculation including first data in COO format indicating multi-dimensional tensor data to the information processing device 100.

[0043] The client device 201 receives the execution result of the CPD calculation for the multi-dimensional tensor data from the information processing device 100. The client device 201 outputs the execution result of the CPD calculation for the multi-dimensional tensor data so that it can be referred to by the system user. The client device 201 is, for example, a server or a PC.

[0044] The information processing device 100 provides a CPD calculation service. The information processing device 100 is used by the system administrator of the CPD calculation system 200. The information processing device 100 receives a request to execute CPD calculation including first data in COO format indicating multi-dimensional tensor data from the client device 201.

[0045] The information processing device 100 converts the first data in COO format indicating the received multi-dimensional tensor data into second data in a new format. The second data is data that can identify a plurality of groups obtained by grouping the attribute information of the elements. The attribute information is a combination of a value and an index for each dimension. The plurality of groups are grouped such that the attribute information including overlapping indexes is included in different groups respectively. The information processing device 100 transmits a request to execute CPD calculation including the second data in the new format to the calculation device 202.

[0046] The information processing device 100 receives the execution result of the CPD calculation for the multi-dimensional tensor data from the calculation device 202. The information processing device 100 transmits the execution result of the CPD calculation for the received multi-dimensional tensor data to the client device 201. The information processing device 100 is, for example, a server or a PC.

[0047] The calculation device 202 receives a request to execute the CPD calculation including the second data in the new format from the information processing device 100. The calculation device 202 executes the CPD calculation based on the second data in the new format. The calculation device 202 transmits the execution result of the CPD calculation to the information processing device 100. The calculation device 202 is, for example, a server or a PC.

[0048] In this way, the CPD calculation system 200 can make the CPD calculation service available on the client device 201 and provide it to the system users. In the following description, mainly the case where the information processing device 100 operates alone will be described.

[0049] (Hardware configuration example of the information processing device 100) Next, a hardware configuration example of the information processing device 100 will be described with reference to FIG. 3.

[0050] FIG. 3 is a block diagram showing a hardware configuration example of the information processing device 100. In FIG. 3, the information processing device 100 includes a CPU (Central Processing Unit) 301, a memory 302, a network I / F (Interface) 303, a recording medium I / F 304, and a recording medium 305. Each component is connected by a bus 300.

[0051] Here, the CPU 301 controls the overall operation of the information processing apparatus 100. The memory 302 includes, for example, a ROM (Read Only Memory), a RAM (Random Access Memory), and a flash ROM. Specifically, for example, the flash ROM and the ROM store various programs, and the RAM is used as the work area of the CPU 301. The programs stored in the memory 302 are loaded into the CPU 301, causing the CPU 301 to execute the encoded processes.

[0052] The network I / F 303 is connected to the network 210 through a communication line and is connected to other computers via the network 210. The network I / F 303 manages the interface between the network 210 and the internal components and controls the input / output of data from / to other computers. The network I / F 303 is, for example, a modem or a LAN adapter.

[0053] The recording medium I / F 304 controls the read / write of data to / from the recording medium 305 according to the control of the CPU 301. The recording medium I / F 304 is, for example, a disk drive, an SSD (Solid State Drive), a USB (Universal Serial Bus) port, etc. The recording medium 305 is a non-volatile memory that stores the data written under the control of the recording medium I / F 304. The recording medium 305 is, for example, a disk, a semiconductor memory, a USB memory, etc. The recording medium 305 may be detachable from the information processing apparatus 100.

[0054] In addition to the components described above, the information processing apparatus 100 may have, for example, a keyboard, a mouse, a display, a printer, a scanner, a microphone, a speaker, etc. Also, the information processing apparatus 100 may have multiple recording medium I / Fs 304 and recording media 305. Further, the information processing apparatus 100 may not have a recording medium I / F 304 or a recording medium 305.

[0055] (Example of the hardware configuration of the client device 201) The hardware configuration example of the client device 201 is the same as the hardware configuration example of the information processing device 100 shown in FIG. 3, and thus the description thereof is omitted.

[0056] (Hardware configuration example of the computing device 202) The hardware configuration example of the computing device 202 is the same as the hardware configuration example of the information processing device 100 shown in FIG. 3, and thus the description thereof is omitted.

[0057] (Functional configuration example of the information processing device 100 according to the first operation example) Next, with reference to FIG. 4, a functional configuration example of the information processing device 100 according to the first operation example, which will be described with reference to FIGS. 5 to 17, will be explained.

[0058] FIG. 4 is a block diagram showing a functional configuration example of the information processing device 100 according to the first operation example. The information processing device 100 includes a storage unit 400, an acquisition unit 401, a first classification unit 402, a second classification unit 403, a calculation unit 404, and an output unit 405.

[0059] The storage unit 400 is realized by a storage area such as the memory 302 and the recording medium 305 shown in FIG. 3, for example. Hereinafter, the case where the storage unit 400 is included in the information processing device 100 will be described, but it is not limited thereto. For example, the storage unit 400 may be included in a device different from the information processing device 100, and the stored content of the storage unit 400 may be referable from the information processing device 100.

[0060] The acquisition unit 401 to the output unit 405 function as an example of a control unit. Specifically, the acquisition unit 401 to the output unit 405 realize their functions by causing the CPU 301 to execute a program stored in a storage area such as the memory 302 and the recording medium 305 shown in FIG. 3, or by the network I / F 303. The processing results of each functional unit are stored in a storage area such as the memory 302 and the recording medium 305 shown in FIG. 3, for example.

[0061] The storage unit 400 stores various information that is referred to or updated in the processing of each functional unit. The storage unit 400 stores, for example, multi-dimensional tensor data. The multi-dimensional tensor data is data that includes a plurality of elements. The elements are associated with the indices of each dimension. The multi-dimensional tensor data may have sparsity. The multi-dimensional tensor data includes, for example, non-zero elements. The multi-dimensional tensor data may include, for example, zero elements. Specifically, the multi-dimensional tensor data is data that enables identification of a combination of the value of each element and the indices of each dimension indicating the position of the element for each element.

[0062] The storage unit 400 stores, for example, first data. The first data is data that enables identification of a combination of the value of each non-zero element included in the multi-dimensional tensor data and the indices of each dimension indicating the position of the element. The first data is, for example, data in COO format. The first data may be, for example, data indicating the result of determining whether each element included in the tensor data is non-zero. Specifically, the first data may be data in which a flag indicating whether each element is non-zero is associated with the index of each dimension of each element. The first data is, for example, acquired by the acquisition unit 401 and stored in the storage unit 400. The first data is, for example, acquired by the first classification unit 402 and stored in the storage unit 400.

[0063] The storage unit 400 stores, for example, second data. The second data is data that enables identification of a plurality of groups obtained by grouping each combination so that combinations with duplicate indices are included in different groups. The second data is, for example, a multi-dimensional array formed by arranging one-dimensional arrays indicating each combination included in the group for each group, and includes a pointer for specifying any one of the combinations included in the group to enable differentiation of the groups. The second data is, for example, generated by the second classification unit 403 and stored in the storage unit 400.

[0064] The acquisition unit 401 acquires various types of information used in the processing of each functional unit. The acquisition unit 401 stores the acquired various types of information in the storage unit 400 or outputs it to each functional unit. Further, the acquisition unit 401 may output the various types of information stored in the storage unit 400 to each functional unit. The acquisition unit 401 acquires various types of information, for example, based on the operation input of the system administrator. The acquisition unit 401 may receive various types of information from a device different from the information processing apparatus 100, for example.

[0065] The acquisition unit 401 acquires tensor data. The acquisition unit 401 acquires tensor data, for example, by receiving tensor data from the client device 201. The acquisition unit 401 acquires tensor data, for example, by accepting the input of tensor data based on the operation input of the system administrator.

[0066] The acquisition unit 401 acquires first data. The acquisition unit 401 acquires first data, for example, by receiving first data from the client device 201. The acquisition unit 401 acquires first data, for example, by accepting the input of first data based on the operation input of the system administrator. When the acquisition unit 401 acquires first data, it does not necessarily have to acquire tensor data.

[0067] The acquisition unit 401 may receive a start trigger for starting the processing of any functional unit. The start trigger is, for example, that there has been a predetermined operation input by the system administrator. The start trigger may be, for example, that predetermined information has been received from another computer. The start trigger may be, for example, that any functional unit has output predetermined information.

[0068] The acquisition unit 401 receives, for example, the acquisition of tensor data as a start trigger for starting the processing of the first classification unit 402, the second classification unit 403, and the calculation unit 404. The acquisition unit 401 receives, for example, the acquisition of the first data as a start trigger for starting the processing of the second classification unit 403 and the calculation unit 404.

[0069] The first classification unit 402 acquires the first data. For example, the first classification unit 402 acquires the first data by generating the first data based on tensor data. Specifically, the first classification unit 402 determines whether each element included in the tensor data is non-zero. For each element determined to be non-zero, the first classification unit 402 acquires the first data by generating the first data that can specify the combination of the value of the element and the indexes of each dimension indicating the position of the element. Thereby, the first classification unit 402 can specify which elements are non-zero and used in the MTTKRP process.

[0070] Specifically, the first classification unit 402 may determine whether each element included in the tensor data is non-zero and acquire the first data indicating the result of determining whether each element included in the tensor data is non-zero. Specifically, for each element included in the tensor data, the first classification unit 402 may acquire the data indicating the result of determining whether the element is non-zero. Thereby, the first classification unit 402 can specify which elements are non-zero and used in the MTTKRP process.

[0071] The second classification unit 403 generates second data based on the acquired first data. For example, the second classification unit 403 groups each combination such that combinations with overlapping indexes are included in different groups respectively. The second classification unit 403 generates, for example, second data that enables identification of a plurality of groups obtained by grouping. Thereby, the second classification unit 403 can enable identification of groups of combinations that can be processed in parallel, and can enable efficient implementation of the MTTKRP process. The second classification unit 403 can reduce the memory usage related to the MTTKRP process.

[0072] Based on the generated second data, the calculation unit 404 performs the MTTKRP process, taking each combination of a plurality of combinations included in a group as an object of parallel processing in the MTTKRP process for tensor data. For example, the calculation unit 404 performs the MTTKRP process with all the combinations included in the group as the object of parallel processing.

[0073] For example, the calculation unit 404 performs the MTTKRP process with a predetermined number of combinations included in the group as the object of parallel processing. Thereby, the calculation unit 404 can efficiently perform the MTTKRP process. The calculation unit 404 can reduce the calculation time. The calculation unit 404 can avoid calculation errors in the MTTKRP process. Further, the calculation unit 404 may perform the ALS process based on the result of performing the MTTKRP process. Thereby, the calculation unit 404 can analyze the tensor data.

[0074] The output unit 405 outputs the processing result of at least any one of the functional units. The output format is, for example, display on a display, print output to a printer, transmission to an external device via the network I / F 303, or storage in a storage area such as the memory 302 or the recording medium 305. Thereby, the output unit 405 can notify the system administrator of the processing result of at least any one of the functional units, and can improve the convenience of the information processing apparatus 100.

[0075] The output unit 405 outputs, for example, the result of performing the MTTKRP process. Thereby, the output unit 405 can facilitate the implementation of ALS. The output unit 405 outputs, for example, the result of performing the ALS process. Thereby, the output unit 405 can make the result of performing the ALS process accessible to the system user and can facilitate the analysis of tensor data.

[0076] Here, the case where the information processing apparatus 100 includes the first classification unit 402, the second classification unit 403, and the calculation unit 404 has been described, but it is not limited thereto. For example, the information processing apparatus 100 may not include the calculation unit 404 and may be communicable with another computer including the calculation unit 404. For example, when the acquisition unit 401 acquires the first data, the information processing apparatus 100 may not include the first classification unit 402.

[0077] (First operation example of the information processing apparatus 100) Next, a first operation example of the information processing apparatus 100 will be described with reference to FIGS. 5 to 17.

[0078] FIGS. 5 to 17 are explanatory diagrams showing a first operation example of the information processing apparatus 100. In FIG. 5, the information processing apparatus 100 acquires first data 500 in COO format corresponding to three-dimensional tensor data. Let each dimension be I, J, and K. The first data 500 can specify the attribute information of each non-zero element included in the three-dimensional tensor data.

[0079] The attribute information includes the value val of the element and the indexes i, j, k of each dimension indicating the position of the element. The first data 500 shows the attribute information of the non-zero elements in each column. In the following description, the attribute information including the value val of the element and the indexes i, j, k of each dimension indicating the position of the element may be denoted as "attribute information (i, j, k, val)".

[0080] The information processing apparatus 100 prepares a bucket 0. The bucket 0 can store the attribute information (i, j, k, val) of the non-zero elements included in the first data 500. The bucket 0 is a storage area for collecting the attribute information (i, j, k, val) of each non-zero element that is included in the first data 500 and whose indices do not overlap with each other.

[0081] The bucket 0 includes an array group 510. The array group 510 includes arrays [] corresponding to each dimension. The array [] corresponding to each dimension has a size of the maximum value of the indices of that dimension + 1. In the example of FIG. 5, since the index values of each dimension are 0 to 3, the size of the array [] corresponding to each dimension is 4.

[0082] The elements of the array [] corresponding to each dimension correspond in order from the beginning to the index values 0, 1, 2, 3 of that dimension. The value of each element of the array [] corresponding to each dimension is a flag indicating whether the attribute information (i, j, k, val) of the non-zero element included in the first data 500 corresponding to the index value of that element has been stored in the bucket 0. If the flag is T (True), it indicates that it has been stored. If the flag is F (False), it indicates that it has not been stored. The value of each element of the array [] corresponding to each dimension is initialized to F. Next, the description of FIG. 6 will be moved on.

[0083] In FIG. 6, the information processing apparatus 100 extracts the attribute information (0, 0, 1, a) of the non-zero element from the first column of the first data 500. The information processing apparatus 100 determines whether the value of the element of the array [] corresponding to the index value of each dimension included in the extracted attribute information (0, 0, 1, a) of the non-zero element in the array group 510 is False.

[0084] Here, if the value of an element of the array [] corresponding to the value of the index of each dimension is one or more True, it is considered that the attribute information of the extracted non-zero element overlaps with the other attribute information already stored in bucket 0. On the other hand, if the value of an element of the array [] corresponding to the value of the index of each dimension is all False, it is considered that the attribute information of the extracted non-zero element does not overlap with the other attribute information even if there is other attribute information already stored in bucket 0. Therefore, it is considered that the attribute information of the extracted non-zero element can be stored in bucket 0.

[0085] In the example of FIG. 6, the information processing apparatus 100 determines that the value of an element of the array [] corresponding to the value of the index of each dimension (0, 0, 1) in the array group 510 is all False. Therefore, the information processing apparatus 100 determines that the attribute information (0, 0, 1, a) of the extracted non-zero element can be stored in bucket 0. Next, the description of FIG. 7 will be moved on.

[0086] In FIG. 7, since the information processing apparatus 100 determines that the attribute information (0, 0, 1, a) of the extracted non-zero element can be stored in bucket 0, the information processing apparatus 100 stores the attribute information (0, 0, 1, a) of the extracted non-zero element in bucket 0. The information processing apparatus 100 sets all the values of the elements of the array [] corresponding to the value of the index of each dimension (0, 0, 1) in the array group 510 to True. Thereby, the information processing apparatus 100 can prevent other attribute information whose index overlaps with the attribute information (0, 0, 1, a) of the non-zero element stored in bucket 0 from being stored. Next, the description of FIG. 8 will be moved on.

[0087] In FIG. 8, the information processing apparatus 100 extracts the attribute information (0, 2, 3, b) of the non-zero element from the second column of the first data 500. The information processing apparatus 100 determines whether the value of an element of the array [] corresponding to the value of each dimension included in the attribute information (0, 2, 3, b) of the extracted non-zero element in the array group 510 is False.

[0088] In the example of FIG. 8, the information processing apparatus 100 determines that the value of the element of the array [0] corresponding to the index value i = 0 of dimension I among the array group 510 is True. Therefore, the information processing apparatus 100 determines that it is not preferable to store the attribute information (0, 2, 3, b) of the extracted non-zero element in bucket 0. Next, the description shifts to FIG. 9.

[0089] In FIG. 9, the information processing apparatus 100 prepares bucket 1. Bucket 1 can store the attribute information (i, j, k, val) of the non-zero elements included in the first data 500. Bucket 1 is a storage area for collecting the attribute information (i, j, k, val) of each non-zero element included in the first data 500 and having non-overlapping indexes with each other.

[0090] Bucket 1 includes an array group 900. The array group 900 is the same as the array group 510. The array group 900 includes arrays [] corresponding to each dimension. The array [] corresponding to each dimension has a size of the maximum value of the index of that dimension + 1. In the example of FIG. 9, since the index values of each dimension are 0 to 3, the size of the array [] corresponding to each dimension is 4.

[0091] The elements of the array [] corresponding to each dimension correspond to the index values 0, 1, 2, 3 of that dimension in order from the beginning. The value of each element of the array [] corresponding to each dimension is a flag indicating whether the attribute information (i, j, k, val) of the non-zero element included in the first data 500 corresponding to the index value corresponding to that element has been stored in bucket 1. The value of each element of the array [] corresponding to each dimension is initialized to F.

[0092] The information processing apparatus 100 determines whether the value of the element of the array [] corresponding to the index value of each dimension included in the attribute information (0, 2, 3, b) of the extracted non-zero element among the array group 900 is False.

[0093] In the example of FIG. 9, the information processing apparatus 100 determines that all the values of the elements of the array [] corresponding to the index values (0, 2, 3) of each dimension in the array group 900 are False. Therefore, the information processing apparatus 100 determines that the attribute information (0, 2, 3, b) of the extracted non-zero elements can be stored in bucket 1. Next, the description of FIG. 10 will be given.

[0094] In FIG. 10, since the information processing apparatus 100 has determined that the attribute information (0, 2, 3, b) of the extracted non-zero elements can be stored in bucket 1, the information processing apparatus 100 stores the attribute information (0, 2, 3, b) of the extracted non-zero elements in bucket 1. The information processing apparatus 100 sets all the values of the elements of the array [] corresponding to the index values (0, 2, 3) of each dimension in the array group 900 to True. Thereby, the information processing apparatus 100 can prevent other attribute information whose index overlaps with the attribute information (0, 2, 3, b) of the non-zero elements stored in bucket 1 from being stored. Next, the description of FIG. 11 will be given.

[0095] In FIG. 11, the information processing apparatus 100 extracts the attribute information (1, 1, 2, c) of non-zero elements from the third column of the first data 500. The information processing apparatus 100 determines whether the values of the elements of the array [] corresponding to the index values of each dimension included in the attribute information (1, 1, 2, c) of the extracted non-zero elements in the array group 510 are False.

[0096] In the example of FIG. 11, the information processing apparatus 100 determines that all the values of the elements of the array [] corresponding to the index values (1, 1, 2) of each dimension in the array group 510 are False. Therefore, the information processing apparatus 100 determines that the attribute information (1, 1, 2, c) of the extracted non-zero elements can be stored in bucket 0. Next, the description of FIG. 12 will be given.

[0097] In FIG. 12, since the information processing apparatus 100 determines that the attribute information (1, 1, 2, c) of the extracted non-zero element can be stored in bucket 0, the information processing apparatus 100 stores the attribute information (1, 1, 2, c) of the extracted non-zero element in bucket 0. The information processing apparatus 100 sets all the values of the elements of the array [] corresponding to the values (1, 1, 2) of the indexes of each dimension in the array group 510 to True. Thereby, the information processing apparatus 100 can prevent other attribute information whose index overlaps with the attribute information (1, 1, 2, c) of the non-zero element stored in bucket 0 from being stored. Next, the description will proceed to FIG. 13.

[0098] In FIG. 13, the information processing apparatus 100 extracts the attribute information (1, 3, 0, d) of the non-zero element from the fourth column of the first data 500. The information processing apparatus 100 determines whether the value of the element of the array [] corresponding to the value of each dimension index included in the attribute information (1, 3, 0, d) of the extracted non-zero element in the array group 510 is False.

[0099] In the example of FIG. 13, the information processing apparatus 100 determines that the value of the element of the array [1] corresponding to the index value i = 1 of dimension I in the array group 510 is True. Therefore, the information processing apparatus 100 determines that it is not preferable to store the attribute information (1, 3, 0, d) of the extracted non-zero element in bucket 0. Next, the description will proceed to FIG. 14.

[0100] In FIG. 14, it is determined whether the value of the element of the array [] corresponding to the value of each dimension index included in the attribute information (1, 3, 0, d) of the extracted non-zero element in the array group 900 is False.

[0101] In the example of FIG. 14, the information processing apparatus 100 determines that all the values of the elements of the array [] corresponding to the values (1, 3, 0) of the indexes of each dimension in the array group 900 are False. Therefore, the information processing apparatus 100 determines that the attribute information (1, 3, 0, d) of the extracted non-zero element can be stored in bucket 1. Next, the description will proceed to FIG. 15.

[0102] In FIG. 15, since the information processing apparatus 100 determines that the attribute information (1, 3, 0, d) of the extracted non-zero element can be stored in bucket 1, the information processing apparatus 100 stores the attribute information (1, 3, 0, d) of the extracted non-zero element in bucket 1. The information processing apparatus 100 sets all the values of the elements of the array [] corresponding to the values (1, 3, 0) of the indexes of each dimension in the array group 900 to True. Thereby, the information processing apparatus 100 can prevent other attribute information whose indexes overlap with the attribute information (1, 3, 0, d) of the non-zero element stored in bucket 1 from being stored. Next, the description will proceed to FIG. 16.

[0103] In FIG. 16, similarly, the information processing apparatus 100 extracts the attribute information of non-zero elements in order from the 5th column and subsequent columns of the first data 500, and determines whether each can be stored in a respective bucket. The information processing apparatus 100 stores the extracted attribute information of non-zero elements in the buckets determined to be storable. As a result, the information processing apparatus 100 stores the attribute information (2, 2, 3, e), the attribute information (3, 3, 0, g), etc. in bucket 0. Also, the information processing apparatus 100 stores the attribute information (2, 0, 1, f), the attribute information (3, 1, 2, h), etc. in bucket 1. Next, the description will proceed to FIG. 17.

[0104] In FIG. 17, the information processing apparatus 100 arranges the attribute information stored in each bucket in order, and generates a pointer array ptr[] that designates the attribute information so that the groups of the attribute information stored in each bucket can be distinguished.

[0105] In the example of FIG. 17, the information processing apparatus 100 prepares a storage destination area 1700. The information processing apparatus 100 sets a pointer indicating the head position 0 of the storage destination area 1700, which is the storage destination of the attribute information included in bucket 0, in the pointer array ptr[0]. The information processing apparatus 100 stores each of the attribute information stored in bucket 0 in order from the position 0 designated by the pointer array ptr[0] in the storage destination area 1700.

[0106] The information processing apparatus 100 is a storage destination for storing the attribute information included in bucket 1, and sets a pointer indicating position 4, which is the next position after the position where the storage of the attribute information included in bucket 0 has ended, to pointer array ptr[1]. The information processing apparatus 100 stores each piece of attribute information stored in bucket 1 in order from position 4 specified by pointer array ptr[1] in storage destination area 1700.

[0107] The information processing apparatus 100 sets a pointer indicating position 8, which is the next position after the position where the storage of the attribute information included in bucket 1 has ended, to pointer array ptr[2]. The information processing apparatus 100 outputs the data formed in storage destination area 1700 as second data to be used in the MTTKRP process. Thereby, the information processing apparatus 100 can identify a group in which non-zero elements that can be processed in parallel in the MTTKRP process are grouped together. For this reason, the information processing apparatus 100 can reduce the calculation time required for the MTTKRP process.

[0108] The information processing apparatus 100 is, for example, in ALS, X (1) (C◎B) calculation, and X (2) (C◎A) calculation, and X (3) (A◎B) calculation, etc., can reduce the calculation time and memory usage required for the MTTKRP process. Therefore, the information processing apparatus 100 can easily eliminate the bottleneck of ALS and can reduce the calculation time required for analyzing tensor data by ALS.

[0109] Also, the information processing apparatus 100 can, for example, avoid conflicts between different operations in the MTTKRP process and can accurately perform the MTTKRP process. Further, the information processing apparatus 100 can make the format of the second data the same as the format of the first data 500 and can suppress a decrease in the convenience of the second data.

[0110] Here, the case where the information processing apparatus 100 acquires the first data 500 in the COO format has been described, but it is not limited to this. For example, the information processing apparatus 100 may acquire data in the CSF format instead of the first data 500 in the COO format, and generate second data used for the MTTKRP process based on the data in the CSF format.

[0111] (First classification processing procedure in the first operation example) Next, with reference to FIG. 18, an example of the first classification processing procedure in the first operation example executed by the information processing apparatus 100 will be described. The first classification processing is realized by, for example, the CPU 301 shown in FIG. 3, a storage area such as the memory 302 and the recording medium 305, and the network I / F 303.

[0112] FIG. 18 is a flowchart showing an example of the first classification processing procedure in the first operation example. In FIG. 18, the information processing apparatus 100 sets itr = 0 (step S1801).

[0113] Next, the information processing apparatus 100 prepares the itr-th new bucket (step S1802). Then, the information processing apparatus 100 determines whether the calculation has been completed for all non-zero elements in the tensor data (step S1803).

[0114] Here, when the calculation has been completed for all non-zero elements (step S1803: Yes), the information processing apparatus 100 ends the first classification processing. On the other hand, when the calculation has not been completed for any non-zero element (step S1803: No), the information processing apparatus 100 proceeds to the process of step S1804.

[0115] In step S1804, the information processing apparatus 100 selects, as operation targets, non-zero elements among the tensor data that have not been selected yet (step S1804). Next, the information processing apparatus 100 determines, based on the index flag, whether other non-zero elements whose indices overlap with the selected non-zero elements of the operation targets are already stored in the itr-th bucket (step S1805).

[0116] Here, if they are already stored (step S1805: Yes), the information processing apparatus 100 proceeds to the process of step S1808. On the other hand, if they are not already stored (step S1805: No), the information processing apparatus 100 proceeds to the process of step S1806.

[0117] In step S1806, the information processing apparatus 100 stores the selected non-zero elements of the operation targets in the itr-th bucket (step S1806). Next, the information processing apparatus 100 sets the flag of the index of the stored non-zero elements to True (step S1807). Then, the information processing apparatus 100 returns to the process of step S1803.

[0118] In step S1808, the information processing apparatus 100 determines whether the (itr + 1)-th bucket exists (step S1808). Here, if the (itr + 1)-th bucket exists (step S1808: Yes), the information processing apparatus 100 proceeds to the process of step S1809. On the other hand, if the (itr + 1)-th bucket does not exist (step S1808: No), the information processing apparatus 100 proceeds to the process of step S1810.

[0119] In step S1809, the information processing apparatus 100 increments itr (step S1809). Then, the information processing apparatus 100 returns to the process of step S1804.

[0120] In step S1810, the information processing apparatus 100 prepares a new itr-th bucket (step S1810). Next, the information processing apparatus 100 stores the non-zero elements of the selected operation target in the prepared itr-th bucket (step S1811).

[0121] Then, the information processing apparatus 100 sets the flag of the index of the stored non-zero elements to True (step S1812). After that, the information processing apparatus 100 returns to the process of step S1803.

[0122] (Second classification processing procedure in the first operation example) Next, with reference to FIG. 19, an example of the second classification processing procedure in the first operation example executed by the information processing apparatus 100 will be described. The second classification processing is realized by, for example, the CPU 301 shown in FIG. 3, a storage area such as the memory 302 and the recording medium 305, and the network I / F 303.

[0123] FIG. 19 is a flowchart showing an example of the second classification processing procedure in the first operation example. In FIG. 19, the information processing apparatus 100 prepares a storage destination (step S1901).

[0124] Next, the information processing apparatus 100 prepares an array ptr[] (step S1902). Then, the information processing apparatus 100 sets itr = 0 (step S1903).

[0125] Next, the information processing apparatus 100 sets ptr[itr] = 0 (step S1904). Then, the information processing apparatus 100 determines whether the processing for all buckets is completed (step S1905).

[0126] Here, when the processing for all buckets is completed (step S1905: Yes), the information processing apparatus 100 ends the second classification processing. On the other hand, when the processing for any bucket is not completed (step S1905: No), the information processing apparatus 100 proceeds to the process of step S1906.

[0127] In step S1906, the information processing apparatus 100 stores each element stored in the itr-th bucket in order from the location specified by ptr[itr] among the storage destinations (step S1906). Next, the information processing apparatus 100 sets ptr[itr + 1] = ptr[itr] + (the number of elements stored in the itr-th bucket) (step S1907).

[0128] Then, the information processing apparatus 100 increments itr (step S1908). After that, the information processing apparatus 100 returns to the process of step S1905.

[0129] (Calculation processing procedure in the first operation example) Next, with reference to FIG. 20, an example of the calculation processing procedure in the first operation example executed by the information processing apparatus 100 will be described. The calculation processing in the first operation example is realized by, for example, the CPU 301 shown in FIG. 3, a storage area such as the memory 302 and the recording medium 305, and the network I / F 303.

[0130] FIG. 20 is a flowchart showing an example of the calculation processing procedure in the first operation example. In FIG. 20, the information processing apparatus 100 sets itr = 0 (step S2001).

[0131] Next, the information processing apparatus 100 determines whether itr is equal to the size of ptr - 1 (step S2002). Here, if itr is equal to the size of ptr - 1 (step S2002: Yes), the information processing apparatus 100 ends the calculation processing. On the other hand, if itr is not equal to the size of ptr - 1 (step S2002: No), the information processing apparatus 100 proceeds to the process of step S2003.

[0132] In step S2003, the information processing apparatus 100 performs a parallel operation on each element (n) stored in the storage area of ptr[itr] to ptr[itr + 1] - 1 in the storage destination (step S2003). The parallel operation is defined by, for example, the following formula (5).

[0133] For f in F: A[I[n], f] += val[n] * C[K[n], f] * B[J[n], f] ···(5)

[0134] Next, the information processing apparatus 100 increments itr (step S2004). Then, the information processing apparatus 100 returns to the process of step S2002. Thereby, the information processing apparatus 100 can perform the MTTKRP process.

[0135] (Functional configuration example of the information processing apparatus 100 according to the second operation example) Next, with reference to FIG. 21, a functional configuration example of the information processing apparatus 100 according to the second operation example described below with reference to FIGS. 22 to 30 will be described.

[0136] FIG. 21 is a block diagram showing a functional configuration example of the information processing apparatus 100 according to the second operation example. The information processing apparatus 100 includes a storage unit 2100, an acquisition unit 2101, a first classification unit 2102, a second classification unit 2103, a calculation unit 2104, and an output unit 2105.

[0137] The storage unit 2100 is realized by a storage area such as the memory 302 and the recording medium 305 shown in FIG. 3, for example. Hereinafter, the case where the storage unit 2100 is included in the information processing apparatus 100 will be described, but it is not limited thereto. For example, the storage unit 2100 may be included in a device different from the information processing apparatus 100, and the stored content of the storage unit 2100 may be referable from the information processing apparatus 100.

[0138] The acquisition unit 2101 to the output unit 2105 function as an example of a control unit. Specifically, the acquisition unit 2101 to the output unit 2105 realize their functions by causing the CPU 301 to execute a program stored in a storage area such as the memory 302 and the recording medium 305 shown in FIG. 3, or by the network I / F 303. The processing results of each functional unit are stored in a storage area such as the memory 302 and the recording medium 305 shown in FIG. 3, for example.

[0139] The storage unit 2100 stores various information that is referenced or updated in the processing of each functional unit. The storage unit 2100 stores, for example, multi-dimensional tensor data. The multi-dimensional tensor data is data that includes a plurality of elements. The elements are associated with the indexes of each dimension. The multi-dimensional tensor data may have sparsity. The multi-dimensional tensor data includes, for example, non-zero elements. The multi-dimensional tensor data may include, for example, zero elements. Specifically, the multi-dimensional tensor data is data that can identify, for each element, the combination of the value of the element and the indexes of each dimension indicating the position of the element.

[0140] The storage unit 2100 stores, for example, first data. The first data is data that can identify, for each non-zero element included in the multi-dimensional tensor data, the combination of the value of the element and the indexes of each dimension indicating the position of the element. The first data is, for example, data in the COO format. Specifically, the first data is data in the SoA (Structure of Array) format. The first data may be, for example, data indicating the result of determining whether each element included in the tensor data is non-zero. Specifically, the first data may be data in which a flag indicating whether each element is non-zero is associated with the indexes of each dimension of each element. The first data is, for example, acquired by the acquisition unit 2101 and stored in the storage unit 2100. The first data is, for example, acquired by the first classification unit 2102 and stored in the storage unit 2100.

[0141] The storage unit 2100 stores, for example, third data. The third data is data that can identify a plurality of groups for a predetermined number of parallel elements. The plurality of groups are obtained by grouping each combination such that combinations with discontinuous indices of the target dimension are not included in the same group according to a predetermined order regarding the indices of the target dimension. The third data includes, for example, for each group, a pointer that designates any one of the combinations included in the group. The third data is generated, for example, by the second classification unit 2103 and stored in the storage unit 2100.

[0142] The acquisition unit 2101 acquires various information used for the processing of each functional unit. The acquisition unit 2101 stores the acquired various information in the storage unit 2100 or outputs it to each functional unit. Also, the acquisition unit 2101 may output the various information stored in the storage unit 2100 to each functional unit. The acquisition unit 2101 acquires various information, for example, based on an operation input by the system administrator. The acquisition unit 2101 may receive various information from a device different from the information processing apparatus 100, for example.

[0143] The acquisition unit 2101 acquires tensor data. The acquisition unit 2101 acquires tensor data, for example, by receiving the tensor data from the client device 201. The acquisition unit 2101 acquires tensor data, for example, by accepting the input of tensor data based on an operation input by the system administrator.

[0144] The acquisition unit 2101 acquires first data. The acquisition unit 2101 acquires first data, for example, by receiving the first data from the client device 201. The acquisition unit 2101 acquires first data, for example, by accepting the input of first data based on an operation input by the system administrator. When the acquisition unit 2101 acquires first data, it does not necessarily have to acquire tensor data.

[0145] The acquisition unit 2101 may receive a start trigger for starting the processing of any functional unit. The start trigger may be, for example, that a predetermined operation input has been made by a system administrator. The start trigger may be, for example, that predetermined information has been received from another computer. The start trigger may be, for example, that any functional unit has output predetermined information.

[0146] The acquisition unit 2101 receives, for example, the acquisition of tensor data as a start trigger for starting the processing of the first classification unit 2102, the second classification unit 2103, and the calculation unit 2104. The acquisition unit 2101 receives, for example, the acquisition of the first data as a start trigger for starting the processing of the second classification unit 2103 and the calculation unit 2104.

[0147] The first classification unit 2102 acquires the first data. The first classification unit 2102 acquires the first data by, for example, generating the first data based on tensor data. Specifically, the first classification unit 2102 determines whether each element included in the tensor data is non-zero. The first classification unit 2102 acquires the first data by generating the first data that can specify the combination of the value of each element determined to be non-zero and the indices of each dimension indicating the position of the element. Thereby, the first classification unit 2102 can specify which elements are non-zero and used in the MTTKRP process.

[0148] Specifically, the first classification unit 2102 may determine whether each element included in the tensor data is non-zero and acquire the first data indicating the result of determining whether each element included in the tensor data is non-zero. Specifically, for each element included in the tensor data, the first classification unit 2102 may acquire the data indicating the result of determining whether the element is non-zero. Thereby, the first classification unit 2102 can specify which elements are non-zero and used in the MTTKRP process.

[0149] The second classification unit 2103 generates third data based on the acquired first data. For example, the second classification unit 2103 groups each combination so that combinations with discontinuous indexes of the target dimension are not included in the same group according to a predetermined order regarding the indexes of the target dimension. The predetermined order is, for example, ascending order or descending order. The second classification unit 2103 generates third data that enables identification of a plurality of groups corresponding to a predetermined number of parallel units obtained by grouping.

[0150] Specifically, the second classification unit 2103 sorts the plurality of combinations so that the indexes of the target dimension are in ascending order. Specifically, the second classification unit 2103 divides the sorted plurality of combinations into a predetermined number of parallel units starting from the beginning. Specifically, the second classification unit 2103 generates third data that enables identification of a plurality of groups with one or more divided combinations as one group each. Thereby, the second classification unit 2103 can identify a plurality of groups that can be processed in parallel, and can efficiently perform the MTTKRP process.

[0151] Based on the generated third data, the calculation unit 2104 performs the MTTKRP process with the plurality of groups as targets for parallel processing in the MTTKRP process regarding the tensor data for the target dimension. For example, the calculation unit 2104 performs an operation regarding each combination included in the group according to a predetermined order, and stores the result of the operation in the temporary area of the group. At this time, for example, each time the operation regarding one or more combinations with the same index of the target dimension included in the group is completed, the calculation unit 2104 reflects the content of the temporary area of the group in the solution matrix.

[0152] As a result, the calculation unit 2104 can efficiently perform the MTTKRP process. The calculation unit 2104 can reduce the calculation time. The calculation unit 2104 can avoid calculation errors in the MTTKRP process. The calculation unit 2104 may further perform the ALS process based on the result of performing the MTTKRP process. Thereby, the calculation unit 2104 can analyze tensor data.

[0153] For example, the calculation unit 2104 converts the first data into the AoS (Array of Structure) format. For example, with reference to the converted first data, the calculation unit 2104 performs the MTTKRP process on a plurality of groups as targets for parallel processing in the MTTKRP process for tensor data with respect to the target dimension based on the third data. Thereby, the calculation unit 2104 can improve the calculation efficiency.

[0154] The output unit 2105 outputs the processing result of at least any one of the functional units. The output format is, for example, display on a display, print output to a printer, transmission to an external device via the network I / F 303, or storage in a storage area such as the memory 302 or the recording medium 305. Thereby, the output unit 2105 can notify the system administrator of the processing result of at least any one of the functional units, and can improve the convenience of the information processing apparatus 100.

[0155] For example, the output unit 2105 outputs the result of performing the MTTKRP process. Thereby, the output unit 2105 can facilitate the implementation of ALS. For example, the output unit 2105 outputs the result of performing the ALS process. Thereby, the output unit 2105 can make the result of performing the ALS process referable by the system user, and can facilitate the analysis of tensor data.

[0156] Here, the case where the information processing apparatus 100 includes the first classification unit 2102, the second classification unit 2103, and the calculation unit 2104 has been described, but it is not limited to this. For example, the information processing apparatus 100 may not include the calculation unit 2104 and may be communicable with another computer including the calculation unit 2104. For example, when the acquisition unit 2101 acquires the first data, the information processing apparatus 100 may not include the first classification unit 2102.

[0157] (Second operation example of information processing apparatus 100) Next, a second operation example of the information processing apparatus 100 will be described with reference to FIGS. 22 to 30.

[0158] FIGS. 22 to 30 are explanatory diagrams showing a second operation example of the information processing apparatus 100. In FIG. 22, the information processing apparatus 100 acquires first data 2200 in COO format corresponding to three-dimensional tensor data. Let each dimension be I, J, and K. The first data 2200 enables identification of the attribute information of each non-zero element included in the three-dimensional tensor data. The first data 2200 is, for example, data in SoA format.

[0159] The attribute information includes the value val of the element and the indices i, j, k of each dimension indicating the position of the element. The first data 2200 indicates the attribute information of non-zero elements in each column. In the following description, the attribute information including the value val of the element and the indices i, j, k of each dimension indicating the position of the element may be denoted as "attribute information (i, j, k, val)".

[0160] The information processing apparatus 100 processes dimension J and prepares an array pair 2210 corresponding to dimension J. The array pair 2210 includes an index array 2211 and a permutation array 2212. The index array 2211 is a one-dimensional array. As the value of each element of the index array 2211, the index j of dimension J of each element of the first data 2200 is set along the order of appearance of each element of the first data 2200. The permutation array 2212 is a one-dimensional array. As the value of each element of the permutation array 2212, integers greater than or equal to 0 are set in ascending order. Next, we will move on to the explanation of FIG. 23.

[0161] In FIG. 23, the information processing apparatus 100 sorts each element of the index array 2211 in ascending order. The information processing apparatus 100 sorts each element of the permutation array 2212 in accordance with the sorting of each element of the index array 2211. For example, when the information processing apparatus 100 moves the i-th element of the index array 2211 to the j-th position, it similarly moves the i-th element of the permutation array 2212 to the j-th position. The information processing apparatus 100 deletes the index array 2211.

[0162] Thereby, the information processing apparatus 100 can make each element of the first data 2200 identifiable in ascending order of the index j of dimension J by the permutation array 2212. Next, we will move on to the explanation of FIG. 24.

[0163] In FIG. 24, assume that the information processing apparatus 100 generates permutation arrays 2401 and 2402 in the same manner as when it processes dimension J, with dimensions I and K as the processing targets. Next, we will move on to the explanation of FIG. 25.

[0164] In FIG. 25, the information processing apparatus 100 converts the first data 2200 into data 2500 in the AoS format. In the first data 2200, each element is arranged in the memory area in the arrangement 2501. In contrast, in the data 2500, each element is arranged in the memory area in the arrangement 2502. Thereby, the information processing apparatus 100 can easily access the attribute information of the elements sequentially, and can improve the access efficiency for each element. Next, the description will proceed to FIG. 26.

[0165] In FIG. 26, the information processing apparatus 100 performs the calculation of X in ALS by parallel processing (2) (C◎A), and prepares a carry array for a predetermined number of parallel processes with dimension J as the processing target. The predetermined number of parallel processes is the number that the information processing apparatus 100 can perform in parallel. In the example of FIG. 26, the predetermined number of parallel processes is 3. The carry array is initialized to 0. The information processing apparatus 100 prepares, for example, a carry array T0, a carry array T1, and a carry array T2.

[0166] The information processing apparatus 100 classifies each element of the permutation array 2212 into a plurality of groups for a predetermined number of parallel processes that are the targets of parallel processing. The information processing apparatus 100 classifies, for example, the first three elements, the subsequent three elements, and the last two elements into one group each. Thereby, the information processing apparatus 100 can identify, for example, the first element, the third element, and the sixth element in the data 2500 specified by the three elements classified into the same group as the processing unit. Next, the description will proceed to FIG. 27.

[0167] In FIG. 27, the information processing apparatus 100 handles each group with a different thread and performs X in ALS (2)Perform the calculation of (C◎A). The thread is executed, for example, by the information processing apparatus 100. The thread may be executed by, for example, something other than the information processing apparatus 100. The information processing apparatus 100 performs, for example, a calculation using three elements in the data 2500 specified by the first three elements at the head of the permutation array 2212 classified into the first group, using the carry array T0 by thread 0.

[0168] The information processing apparatus 100 selects, for example, the first element of the data 2500 based on the first element classified into the first group by thread 0. The information processing apparatus 100 is, for example, X (2) Store the result of performing the calculation for the selected first element in the calculation of (C◎A) in the carry array T0. The information processing apparatus 100 stores, for example, the index j of the dimension J of the first element as prev.

[0169] The information processing apparatus 100 selects, for example, the sixth element of the data 2500 based on the second element classified into the first group by thread 0. The information processing apparatus 100 determines, for example, whether the index j of the dimension J of the selected sixth element matches prev. Since the information processing apparatus 100 determines that they match, for example, X (2) Add and store the result of performing the calculation for the selected sixth element in the calculation of (C◎A) in the carry array T0. The information processing apparatus 100 stores, for example, the index j of the dimension J of the selected sixth element as prev. Next, proceed to the description of FIG. 28.

[0170] In FIG. 28, for example, the information processing apparatus 100 selects the third element of the data 2500 based on the third element classified into the first group by thread 0. The information processing apparatus 100 determines, for example, whether the index j of dimension J of the selected third element matches prev. Since the information processing apparatus 100 determines that they do not match, for example, it writes and stores the content of the carry array T0 in the output matrix B. Then, after initializing the carry array T0, for example, the information processing apparatus 100 (2) Stores the result of the calculation for the selected third element in the carry array T0 in the calculation of (C◎A). The information processing apparatus 100 stores, for example, the index j of dimension J of the selected third element as prev. Next, the description proceeds to FIG. 29.

[0171] In FIG. 29, assume that the information processing apparatus 100, for example, performs calculations using three elements in the data 2500 with the carry array T1 by thread 1 in parallel with thread 0, similar to thread 0. Also, assume that the information processing apparatus 100, for example, performs calculations using three elements in the data 2500 with the carry array T2 by thread 2 in parallel with thread 0, similar to thread 0.

[0172] When the calculations by thread 0, thread 1, and thread 2 are completed, the information processing apparatus 100 writes and stores the content of the carry array T0, the carry array T1, and the carry array T2 in the output matrix B in order. Thereby, the information processing apparatus 100 (2) Can perform the calculation of X(C◎A) in ALS by parallel processing. Similarly, the information processing apparatus 100 (1) Can perform the calculation of X(C◎B) in ALS and (3) Can perform the calculation of X(A◎B) in ALS. Therefore, the information processing apparatus 100 can reduce the calculation time required for the MTTKRP process.

[0173] The information processing apparatus 100, for example, in ALS,(1) (Calculation of (C◎B) and X (2) (Calculation of (C◎A) and X (3) It is possible to reduce the calculation time and memory usage for MTTKRP processing such as the calculation of (A◎B). Therefore, the information processing apparatus 100 can easily eliminate the bottleneck of ALS, and can reduce the calculation time required for analyzing tensor data by ALS.

[0174] In addition, the information processing apparatus 100 can avoid conflicts between different operations in MTTKRP processing based on, for example, a sorted permutation array, and can accurately perform MTTKRP processing. Further, the information processing apparatus 100 can equalize the processing amounts of the respective threads. For this reason, the information processing apparatus 100 can efficiently perform the calculation of X (2) (C◎A) in ALS. The information processing apparatus 100 can easily reduce the calculation time required for MTTKRP processing regardless of the number of elements with duplicate indexes in the first data 2200. Next, we will move on to the description of FIG. 30.

[0175] In FIG. 30, the calculation efficiency when the information processing apparatus 100 performs ALS is shown. The graph 3000 in FIG. 30 shows the magnification NEW with respect to the calculation efficiency when ALS is performed on the data in the COO format by the conventional method as a reference for the calculation efficiency when the information processing apparatus 100 performs ALS. Further, the graph 3000 in FIG. 30 shows the magnification with respect to the calculation efficiency when ALS is performed on the data in the CSF format by the conventional method, with respect to the calculation efficiency when ALS is performed on the data in the COO format by the conventional method as a reference. Further, the graph 3000 in FIG. 30 shows the magnification with respect to the calculation efficiency when ALS is performed in HiCOO, with respect to the calculation efficiency when ALS is performed on the data in the COO format by the conventional method as a reference.

[0176] As shown in graph 3000 of FIG. 30, information processing apparatus 100 can improve the calculation efficiency of ALS. Further, information processing apparatus 100 can more easily improve the calculation efficiency of ALS based on tensor data having a relatively large number of dimensions or a relatively large number of non-zero elements as compared with conventional methods.

[0177] (Preparation processing procedure) Next, with reference to FIG. 31, an example of a preparation processing procedure executed by information processing apparatus 100 will be described. The preparation processing is realized by, for example, CPU 301 shown in FIG. 3, a storage area such as memory 302 and recording medium 305, and network I / F 303.

[0178] FIG. 31 is a flowchart showing an example of a preparation processing procedure. In FIG. 31, information processing apparatus 100 determines whether or not all dimensions have been selected as processing targets (step S3101). Here, if all dimensions have been selected (step S3101: Yes), information processing apparatus 100 proceeds to the processing of step S3107. On the other hand, if there are still dimensions that have not been selected (step S3101: No), information processing apparatus 100 proceeds to the processing of step S3102.

[0179] In step S3102, information processing apparatus 100 selects any one of the dimensions that have not yet been selected as a processing target (step S3102). Next, information processing apparatus 100 copies an index array in which the indexes of the selected dimension are arranged in the order of appearance on the data in COO format (step S3103). Then, information processing apparatus 100 initializes a permutation array (step S3104).

[0180] Next, information processing apparatus 100 associates the copy of the index array with the permutation array and sorts the permutation array based on the copy of the index array (step S3105). Then, information processing apparatus 100 deletes the copy of the index array (step S3106). Thereafter, information processing apparatus 100 returns to the processing of step S3101.

[0181] In step S3107, the information processing apparatus 100 secures an area for storing data in AoS format corresponding to the tensor data (step S3107). Next, the information processing apparatus 100 converts the data in SoA format corresponding to the tensor data into data in AoS format (step S3108). Then, the information processing apparatus 100 deletes the data in SoA format corresponding to the tensor data (step S3109). After that, the information processing apparatus 100 ends the preparation process.

[0182] (Calculation processing procedure) Next, with reference to FIG. 32, an example of the calculation processing procedure executed by the information processing apparatus 100 will be described. The calculation processing is realized by, for example, the CPU 301 shown in FIG. 3, a storage area such as the memory 302 and the recording medium 305, and the network I / F 303.

[0183] FIG. 32 is a flowchart showing an example of the calculation processing procedure. In FIG. 32, the information processing apparatus 100 secures a plurality of carry areas according to the number of parallel operations (step S3201).

[0184] Next, the information processing apparatus 100 initializes the plurality of carry areas (step S3202). Then, the information processing apparatus 100 executes the parallel calculation processing described later with reference to FIG. 33 (step S3203).

[0185] Next, the information processing apparatus 100 executes the sequential calculation processing described later with reference to FIG. 34 (step S3204). Then, the information processing apparatus 100 releases the plurality of carry areas (step S3205). After that, the information processing apparatus 100 ends the calculation processing.

[0186] (Parallel calculation processing procedure) Next, with reference to FIG. 33, an example of the parallel calculation processing procedure executed by the information processing apparatus 100 will be described. The parallel calculation processing is realized by, for example, the CPU 301 shown in FIG. 3, a storage area such as the memory 302 and the recording medium 305, and the network I / F 303.

[0187] The parallel calculation processing shown in FIG. 33 corresponds to the calculation of X (2) (C◎A). For the parallel calculation processing corresponding to the calculation of X (1) (C◎B) and X (3) (A◎B), since it is the same as the parallel calculation processing shown in FIG. 33, the description thereof will be omitted.

[0188] FIG. 33 is a flowchart showing an example of the parallel calculation processing procedure. In FIG. 33, the information processing apparatus 100 designates the scope of responsibility of each of a plurality of threads corresponding to the number of threads based on the sorted permutation array (step S3301).

[0189] Next, the information processing apparatus 100 sets prev = X[permutation_J[start]].J for each thread (step S3302). prev is set to, for example, the index of the non-zero element accessed immediately before or first. X[].J indicates the index j of dimension J of any element of the data in AoS format. permutation_J[start] indicates the first element of the permutation array.

[0190] Then, the information processing apparatus 100 determines whether parallel operations have been performed on all non-zero elements included in the scope of responsibility of each thread (step S3303). Here, when parallel operations have been performed on all non-zero elements (step S3303: Yes), the information processing apparatus 100 ends the parallel calculation processing. On the other hand, when parallel operations have not been performed on any non-zero element (step S3303: No), the information processing apparatus 100 proceeds to the processing of step S3304.

[0191] In step S3304, the information processing apparatus 100 determines whether prev == X[n].J for the nth non-zero element (n) included in the responsibility range of at least any one thread (step S3304). Here, when prev == X[n].J (step S3304: Yes), the information processing apparatus 100 proceeds to the process of step S3307. On the other hand, when prev != X[n].J (step S3304: No), the information processing apparatus 100 selects the thread for which it has determined that prev != X[n].J and proceeds to the process of step S3305.

[0192] In step S3305, the information processing apparatus 100 reads the content of the carry area corresponding to the selected thread, writes it to the output matrix B, and initializes the carry area (step S3305). Next, the information processing apparatus 100 sets prev = X[n].J for the selected thread (step S3306). Then, the information processing apparatus 100 proceeds to the process of step S3307.

[0193] In step S3307, the information processing apparatus 100 performs a parallel operation (step S3307). The information processing apparatus 100 performs a parallel operation, for example, by causing each thread to perform a predetermined operation using the nth non-zero element (n) included in the responsibility range of the thread in the carry area corresponding to each thread. The parallel operation is defined, for example, by the following formula (6). Then, the information processing apparatus 100 returns to the process of step S3303.

[0194] for f in F : B[X[n].J,f]+=X[n].val*C[X[n].K,f]*A[X[n].I,f] ···(6)

[0195] (Sequential calculation processing procedure) Next, with reference to FIG. 34, an example of the sequential calculation processing procedure executed by the information processing apparatus 100 will be described. The sequential calculation processing is realized by, for example, the CPU 301 shown in FIG. 3, storage areas such as the memory 302 and the recording medium 305, and the network I / F 303.

[0196] The sequential calculation processing shown in FIG. 34 corresponds to the calculation of X (2) (C◎A). For the sequential calculation processing corresponding to the calculation of X (1) (C◎B) and X (3) (A◎B), since it is the same as the sequential calculation processing shown in FIG. 34, the description thereof will be omitted.

[0197] FIG. 34 is a flowchart showing an example of the sequential calculation processing procedure. In FIG. 34, the information processing apparatus 100 determines whether or not the sequential operation has been completed for all carry areas (step S3401). Here, if the sequential operation has been completed for all carry areas (step S3401: Yes), the information processing apparatus 100 ends the sequential calculation processing. On the other hand, if the sequential operation has not been completed for any of the carry areas (step S3401: No), the information processing apparatus 100 proceeds to the processing of step S3402.

[0198] In step S3402, the information processing apparatus 100 selects any carry area for which the sequential operation has not yet been completed (step S3402). Next, the information processing apparatus 100 identifies the index J(j) of the non-zero element that the thread in charge of the selected carry area last calculated (step S3403).

[0199] Then, the information processing apparatus 100 performs a sequential operation by causing a predetermined operation to be performed by each thread on the carry area corresponding to each thread (step S3404). The sequential calculation is defined, for example, by the following formula (7). tid is the ID of the thread.

[0200] for f in F : B[r][f]+=carry[tid][f] ···(7)

[0201] As described above, according to the information processing apparatus 100, for each non-zero element included in the multi-dimensional tensor data, it is possible to acquire first data that specifies a combination of the value of the element and the indices of each dimension indicating the position of the element. According to the information processing apparatus 100, based on the first data, it is possible to generate second data that specifies a plurality of groups obtained by grouping each combination so that combinations with overlapping indices are included in different groups. According to the information processing apparatus 100, based on the generated second data, each combination of the plurality of combinations included in the group can be subjected to parallel processing in the MTTKRP process for the tensor data, and the MTTKRP process can be performed. As a result, the information processing apparatus 100 can perform parallel processing of different operations while avoiding competition between different operations for each dimension in the MTTKRP process. Therefore, the information processing apparatus 100 can reduce the calculation time and memory usage amount required for the MTTKRP process.

[0202] According to the information processing apparatus 100, by generating the first data based on the tensor data, the first data can be acquired. As a result, the information processing apparatus 100 can eliminate the need for the system user or system administrator to generate the first data. The information processing apparatus 100 can reduce the work load on the system user or system administrator.

[0203] According to the information processing apparatus 100, for each group, a multi-dimensional array formed by arranging one-dimensional arrays each indicating a combination included in the group can be generated as second data. According to the information processing apparatus 100, a pointer for specifying any one of the combinations included in the group can be included in the second data so as to be able to distinguish the groups. Thereby, the information processing apparatus 100 can generate the second data in the same format as the first data, and can improve the convenience of the second data.

[0204] According to the information processing apparatus 100, tensor data can be acquired. According to the information processing apparatus 100, it can be determined whether each element included in the tensor data is non-zero. According to the information processing apparatus 100, second data can be generated based on the determined result. Thereby, the information processing apparatus 100 can manage without leaving the first data in the storage area, and can reduce the memory usage.

[0205] Note that the information processing method described in the present embodiment can be realized by executing a program prepared in advance on a computer such as a PC or a workstation. The information processing program described in the present embodiment is recorded on a computer-readable recording medium and is executed by being read from the recording medium by a computer. The recording medium is a hard disk, a flexible disk, a CD (Compact Disc)-ROM, an MO (Magneto Optical disc), a DVD (Digital Versatile Disc), or the like. Further, the information processing program described in the present embodiment may be distributed via a network such as the Internet.

[0206] Regarding the above-described embodiment, the following additional remarks are disclosed.

[0207] (Appendix 1) For each non-zero element included in the multi-dimensional tensor data, obtain first data that enables identification of a combination of the value of the element and the indices of each dimension indicating the position of the element, Based on the obtained first data, generate second data that enables identification of a plurality of groups obtained by grouping each of the combinations such that the combinations with overlapping indices are included in different groups, Based on the generated second data, perform the MTTKRP (Matricized Tensor Times Khatri-Rao Product) process on the tensor data by using each combination included in the group as an object of parallel processing in the MTTKRP process, An information processing program characterized by causing a computer to execute the process.

[0208] (Appendix 2) The obtaining process is Based on the tensor data, obtain the first data by generating the first data. The information processing program according to Appendix 1, characterized by this.

[0209] (Appendix 3) The second data is a multi-dimensional array formed by arranging one-dimensional arrays indicating each of the combinations included in the group for each group, and includes a pointer for specifying any one of the combinations included in the group so as to be able to distinguish the groups. The information processing program according to Appendix 1 or 2, characterized by this.

[0210] (Appendix 4) Obtain the tensor data, Determine whether each element included in the tensor data is non-zero, Cause the computer to execute the process, The generating process is Generate the second data based on the determined result. The information processing program according to any one of Appendices 1 to 3, characterized by this.

[0211] (Appendix 5) Based on the obtained first data, generate third data that can identify a plurality of groups of a predetermined number of parallel columns, such that combinations in which the index of the target dimension is discontinuous are not included in the same group according to a predetermined order regarding the index of the target dimension. Based on the generated third data, use the plurality of groups as targets for parallel processing in the MTTKRP process for the tensor data with respect to the target dimension, perform operations on each combination included in the group according to the predetermined order, store the results of the operations in a temporary area of the group, and each time the operation on one or more combinations in the group with the same index of the target dimension is completed, reflect the content of the temporary area of the group in the solution matrix to implement the MTTKRP process. A computer-executable information processing program according to any one of Appendices 1 to 4.

[0212] (Appendix 6) The predetermined order is ascending or descending order of the index of the target dimension. An information processing program according to Appendix 5, characterized by this.

[0213] (Appendix 7) The process to be implemented is Store the combination in the form of an Array of Structure and implement the MTTKRP process. An information processing program according to Appendix 5 or 6, characterized by this.

[0214] (Appendix 8) For each non-zero element included in the multi-dimensional tensor data, obtain first data that can identify the combination of the value of the element and the index of each dimension indicating the position of the element. Based on the obtained first data, generate second data that enables identification of a plurality of groups obtained by grouping each of the combinations such that the combinations with overlapping indexes are included in different groups respectively. Based on the generated second data, perform the MTTKRP (Matricized Tensor Times Khatri-Rao Product) process on the tensor data by using each combination of the multiple combinations included in the group as an object of parallel processing in the MTTKRP process. An information processing method, characterized in that a computer executes the process.

[0215] (Appendix 9) For each non-zero element included in the multi-dimensional tensor data, obtain first data that enables identification of a combination of the value of the element and the indexes of each dimension indicating the position of the element. Based on the obtained first data, generate second data that enables identification of a plurality of groups obtained by grouping each of the combinations such that the combinations with overlapping indexes are included in different groups respectively. Based on the generated second data, perform the MTTKRP (Matricized Tensor Times Khatri-Rao Product) process on the tensor data by using each combination of the multiple combinations included in the group as an object of parallel processing in the MTTKRP process. An information processing apparatus, characterized by having a control unit.

[0216] (Appendix 10) For each non-zero element included in the multi-dimensional tensor data, obtain first data that enables identification of a combination of the value of the element and the indexes of each dimension indicating the position of the element. Based on the obtained first data, generate second data that enables identification of a plurality of groups for a predetermined number of parallel processes, obtained by grouping each of the combinations such that combinations in which the index of the target dimension is discontinuous are not included in the same group, according to a predetermined order regarding the index of the target dimension. Based on the generated second data, for each combination included in the group, perform an operation regarding the combination in accordance with the predetermined order, with the plurality of groups being targets of parallel processing in MTTKRP (Matricized Tensor Times Khatri-Rao Product) processing regarding the tensor data for the target dimension, and store the result of the operation in a temporary area of the group. Each time an operation regarding one or more combinations in which the index of the target dimension included in the group is the same is completed, reflect the content of the temporary area of the group in a matrix to perform the MTTKRP processing. An information processing program characterized by causing a computer to execute the processing.

[0217] (Appendix 11) Obtain first data that enables identification of a combination of the value of each non-zero element included in multi-dimensional tensor data and the index of each dimension indicating the position of the element. Based on the obtained first data, generate second data that enables identification of a plurality of groups for a predetermined number of parallel processes, obtained by grouping each of the combinations such that combinations in which the index of the target dimension is discontinuous are not included in the same group, according to a predetermined order regarding the index of the target dimension. Based on the generated second data, the plurality of groups are used as targets for parallel processing in MTTKRP (Matricized Tensor Times Khatri-Rao Product) processing related to the tensor data for the dimension of the target. According to the predetermined order, operations are performed on each combination of the plurality of combinations included in the group, and the results of the operations are stored in the temporary area of the group. Each time the operation on one or more combinations in the group with the same index of the dimension of the target is completed, the content of the temporary area of the group is reflected in the matrix to perform the MTTKRP processing. An information processing method characterized in that a computer executes the processing.

[0218] (Appendix 12) For each non-zero element included in the multi-dimensional tensor data, first data that enables identification of the combination of the value of the element and the indexes of each dimension indicating the position of the element is acquired. Based on the acquired first data, according to a predetermined order regarding the indexes of the dimension of the target, second data is generated that enables identification of a plurality of groups for a predetermined number of parallelisms, obtained by grouping each of the combinations so that combinations with discontinuous indexes of the dimension of the target are not included in the same group. Based on the generated second data, the plurality of groups are used as targets for parallel processing in MTTKRP (Matricized Tensor Times Khatri-Rao Product) processing related to the tensor data for the dimension of the target. According to the predetermined order, operations are performed on each combination of the plurality of combinations included in the group, and the results of the operations are stored in the temporary area of the group. Each time the operation on one or more combinations in the group with the same index of the dimension of the target is completed, the content of the temporary area of the group is reflected in the matrix to perform the MTTKRP processing. An information processing apparatus characterized by having a control unit.

Explanation of Signs

[0219] 100 Information processing device 101,500,2200 First data 102 Second data 110 Code 200 CPD calculation system 201 Client device 202 Calculation device 210 Network 300 Bus 301 CPU 302 Memory 303 Network I / F 304 Recording medium I / F 305 Recording medium 400,2100 Storage unit 401,2101 Acquisition unit 402,2102 First classification unit 403,2103 Second classification unit 404,2104 Calculation unit 405,2105 Output unit 510,900 Array group 1700 Storage destination area 2210 Array pair 2211 Index array 2212,2401,2402 Permutation array 2500 Data 2501,2502 Arrangement 3000 Graph

Claims

1. For each non-zero element included in the multi-dimensional tensor data, obtain first data that enables identification of a combination of the value of the element and the indices of each dimension indicating the position of the element, Based on the obtained first data, generate second data that enables identification of a plurality of groups obtained by grouping each of the combinations such that the combinations with overlapping indices are included in different groups, Based on the generated second data, perform the MTTKR P (Matricized Tensor Times Khatri-Rao Product) process on the tensor data by using each of the combinations included in the group as an object of parallel processing in the MTTKR P process, An information processing program characterized by causing a computer to execute the process.

2. The obtaining process obtains the first data by generating the first data based on the tensor data, The information processing program according to claim 1, characterized in that.

3. The second data is a multi-dimensional array formed by arranging one-dimensional arrays indicating each of the combinations included in the group for each group, and includes a pointer for specifying any one of the combinations included in the group so as to be able to distinguish the groups. The information processing program according to claim 1 or 2, characterized in that.

4. Based on the obtained first data, generate third data that enables identification of a plurality of groups for a predetermined number of parallel processes obtained by grouping each of the combinations such that the combinations with discontinuous indices of the target dimension are not included in the same group according to a predetermined order regarding the indices of the target dimension, Based on the generated third data, the plurality of groups are used as targets for parallel processing in the MTTKRP processing related to the tensor data for the dimension of the target, and operations are performed on each combination included in the group in the predetermined order, and the results of the operations are stored in the temporary area of the group. Each time the operation on one or more combinations with the same index of the dimension of the target included in the group is completed, the content of the temporary area of the group is reflected in the solution matrix to implement the MTTKRP processing. An information processing program according to any one of claims 1 to 3, characterized in that the computer is caused to execute the processing.

5. The information processing program according to claim 4, characterized in that the predetermined order is ascending or descending order of the index of the dimension of the target.

6. The processing to be implemented is storing the combination in the form of an Array of Structure and implementing the MTTKRP processing, and the information processing program according to claim 4 or 5, characterized in that.

7. For each non-zero element included in the multi-dimensional tensor data, obtain first data that can specify a combination of the value of the element and the indices of each dimension indicating the position of the element, Based on the obtained first data, generate second data that can specify a plurality of groups obtained by grouping each of the combinations so that the combinations with overlapping indices are included in different groups, Based on the generated second data, each combination of the plurality of combinations included in the group is used as a target for parallel processing in the MTTKRP (Matricized Tensor Times Khatri-Rao Product) processing related to the tensor data, and the MTTKRP processing is implemented. An information processing method, characterized in that a computer executes the processing.

8. For each non-zero element included in the multi-dimensional tensor data, obtain first data that can specify a combination of the value of the element and the indices of each dimension indicating the position of the element, Based on the obtained first data, generate second data that enables identification of a plurality of groups obtained by grouping each of the combinations such that the combinations with overlapping indexes are included in different groups respectively. Based on the generated second data, perform the MTTKRP (Matricized Tensor Times Khatri-Rao Product) process on each combination of the plurality of combinations included in the group as an object of parallel processing in the MTTKRP process related to the tensor data. An information processing apparatus characterized by having a control unit.

9. For each non-zero element included in the multi-dimensional tensor data, obtain first data that enables identification of a combination of the value of the element and indexes of each dimension indicating the position of the element. Based on the obtained first data, generate second data that enables identification of a plurality of groups corresponding to a predetermined number of parallelisms, obtained by grouping each of the combinations such that combinations with discontinuous indexes in the target dimension do not belong to the same group according to a predetermined order regarding the indexes of the target dimension. Based on the generated second data, perform operations on each combination included in the group regarding the MTTKRP (Matricized Tensor Times Khatri-Rao Product) process related to the tensor data for the target dimension in accordance with the predetermined order, store the results of the operations in a temporary area of the group, and each time the operation on one or more combinations with the same index in the target dimension included in the group is completed, reflect the content of the temporary area of the group in the matrix decomposition to perform the MTTKRP process. An information processing program characterized by causing a computer to execute the process.

10. For each non-zero element included in the multi-dimensional tensor data, obtain first data that enables identification of a combination of the value of the element and indexes of each dimension indicating the position of the element. Based on the obtained first data, generate second data that enables identification of a plurality of groups for a predetermined number of parallel columns, where each combination is grouped such that combinations with discontinuous indices of the target dimension are not included in the same group, according to a predetermined order regarding the indices of the target dimension. Based on the generated second data, use the plurality of groups as targets for parallel processing in the MTTKR P (Matricized Tensor Times Khatri-Rao Product) process for the tensor data with respect to the target dimension. Perform operations on each combination included in the groups according to the predetermined order, store the results of the operations in a temporary area of the group, and each time the operation on one or more combinations with the same index of the target dimension included in the group is completed, reflect the content of the temporary area of the group in the matrix to implement the MTTKR P process. An information processing method characterized in that a computer executes the process. [

11. ] For each non-zero element included in multi-dimensional tensor data, obtain first data that enables identification of a combination of the value of the element and the indices of each dimension indicating the position of the element. Based on the obtained first data, generate second data that enables identification of a plurality of groups for a predetermined number of parallel columns, where each combination is grouped such that combinations with discontinuous indices of the target dimension are not included in the same group, according to a predetermined order regarding the indices of the target dimension. Based on the generated second data, the plurality of groups are used as targets for parallel processing in the MTTKR P (Matricized Tensor Times Khatri-Rao Product) process related to the tensor data for the dimension of the target, and operations are performed on each combination among the plurality of combinations included in the group in the predetermined order, and the results of the operations are stored in the temporary area of the group. Each time the operation on one or more combinations in which the indexes of the dimension of the target included in the group are the same is completed, the MTTKR P process is implemented by reflecting the content of the temporary area of the group in the unfolding matrix. An information processing apparatus characterized by having a control unit.

Citation Information

Patent Citations

  • A system, apparatus, and method for expanding a memory source into a destination register and compressing the source register into a destination memory location.

    JP2014513341A

  • Computer-implemented system and method for efficient sparse matrix representation and processing

    JP2016119084A

  • Tensor data calculation device, tensor data calculation method, and program

    JP2016139391A

  • Tensor factor decomposition processing apparatus, tensor factor decomposition processing method and tensor factor decomposition processing program

    JP2018128708A

  • Matrix arithmetic device, matrix arithmetic method, and matrix arithmetic program

    JP2019148969A