ALS-based parallel tensor filling method and system

By adopting ALS-based parallel tensor filling method in GPU parallel computing, the problems of computational load imbalance and large global memory access overhead are solved, and efficient tensor filling and computational performance improvements are achieved.

CN120216853AActive Publication Date: 2025-06-27HUNAN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510704237.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-06-27
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The existing technology has problems of unbalanced computing load and large global memory access overhead in GPU parallel computing, which limits the performance improvement of tensor filling algorithms on large-scale data.

Method used

A parallel tensor filling method based on ALS is adopted to create tensors by collecting data feature information, generate initial factor matrix, and divide molecular tensors according to the performance of hardware equipment for parallel processing. The sub-tensors are stored using the COO format and the fixed factor matrix is ​​placed in shared memory during processing to speed up access, alternating update of the factor matrix to achieve missing value padding.

Benefits of technology

It significantly improves parallel computing efficiency, reduces memory access overhead, speeds up convergence speed, is suitable for efficient filling of large-scale sparse tensors, and achieves the improvement of load balancing and data access efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216853A_ABST
    Figure CN120216853A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric digital data processing, and discloses an ALS-based parallel tensor filling method and system, and the method comprises the steps: collecting feature information of a data set needing to be filled, building a tensor according to a preset data feature, and generating an initial factor matrix of the tensor; obtaining the number and position of tensor effective values, dividing sub-tensors according to hardware equipment performance to obtain a plurality of sub-tensors processed in parallel, and storing the sub-tensors in a COO format; a processing unit is distributed to one sub-tensor, reduction operation is carried out, and updated values of each row are assigned to the currently updated factor matrix; the three factor matrixes are alternately updated and iterated for multiple rounds, after each round of iteration is completed, the updated factor matrixes are subjected to outer product to obtain filled tensors, the loss between the front tensor and the rear tensor is calculated, the factor matrixes are alternately optimized, and missing value filling is achieved; the parallel processing efficiency is improved, and the overall calculation delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and particularly relates to a parallel tensor filling method and system based on ALS. Background Art

[0002] A tensor is a multi-dimensional array structure in mathematics, which can be regarded as a high-dimensional generalization of vectors and matrices. Tensor filling refers to the process of estimating and filling these missing values through certain algorithms in the case of partial data loss, aiming to restore the integrity and availability of the data, and is mainly used to ensure that data of different sizes have a unified shape in deep learning, image processing, and data preprocessing, so as to facilitate batch processing and model training. Network traffic monitoring is a key task in modern network management and security, including anomaly detection, traffic prediction, and attack identification, etc. In recent years, due to its advantages in high-dimensional data recovery and feature extraction, tensor filling technology has gradually been introduced into the network traffic monitoring scenario.

[0003] Due to the limitations of data acquisition devices and the complexity of the network environment, the obtained traffic data often has missing or incomplete situations. Network traffic data can be modeled as a multi-dimensional tensor (such as [time × source IP × destination IP × port × protocol]), and the missing part is inferred through the low-rank tensor filling method to restore the traffic pattern. In practical applications, due to the limitations of data acquisition devices and the complexity of the network environment, the obtained network traffic data often has a large number of missing or incomplete situations. However, with the continuous growth of the network scale, processing large-scale traffic tensors faces a computing bottleneck, and GPU acceleration has become an important means to improve computing efficiency.

[0004] Most of the existing methods mainly optimize the storage format, such as using sparse matrix storage, and do not introduce dynamic task allocation, which may lead to unbalanced computing loads and affect the overall parallel efficiency; although some methods utilize GPU acceleration, the computing process mainly relies on global memory access, without considering that accessing shared memory is 100 times faster than accessing global memory, so the access to GPU memory is not optimized, resulting in a large access overhead, and there is also great room for improvement in computing performance here.

[0005] Based on the above analysis, although the existing methods have achieved certain optimizations in GPU parallel computing, there are still problems such as unbalanced computing loads and large global memory access overheads, which limit the performance improvement of tensor filling algorithms on large-scale data. Summary of the Invention

[0006] The purpose of the present invention is to solve the above problems and design a parallel tensor filling method and system based on ALS.

[0007] The first aspect of the present invention provides a parallel tensor filling method based on ALS, and the method includes the following steps: Collect the feature information of the dataset to be filled, and establish a tensor according to the preset data features to generate the initial factor matrix of the tensor; Obtain the number and positions of the valid values of the tensor, divide the sub-tensors according to the performance of the hardware device to obtain multiple sub-tensors for parallel processing, and store the sub-tensors in the COO format; Allocate a processing unit for a sub-tensor and perform a reduction operation, and assign the updated value of each row to the factor matrix being updated currently; Alternately update the three factor matrices and perform multiple rounds of iteration. After each round of iteration, perform an outer product on the updated factor matrices to obtain the filled tensor, calculate the loss between the front and back tensors, and alternately optimize each factor matrix to achieve missing value filling.

[0008] Optionally, in the first implementation manner of the first aspect of the present invention, obtain three features in the dataset to be filled that have real physical relationships with each other, establish a corresponding three-dimensional tensor, fill the missing values with 0 values, and perform feature scaling using the maximum-minimum normalization to obtain the normalized tensor.

[0009] Optionally, in the second implementation manner of the first aspect of the present invention, realize the load migration of the computing task by dynamically adjusting the sub-tensor boundary. Use the valid values of the sub-tensor as the computing load index, expand the adjacent slices of the high-load sub-tensor to the low-load area, and reallocate the valid values by adjusting the slice thickness so that the load of the adjusted sub-tensor satisfies: ; where is the tolerance parameter, is the optimal number of valid values for each sub-tensor, is the effective load value of the sub-tensor.

[0010] Optionally, in the third implementation manner of the first aspect of the present invention, traverse the entire factor matrix to find all non-zero elements. For each non-zero element, record its row index, column index, and the corresponding value, and store these indexes and values in three arrays respectively.

[0011] Optionally, in the fourth implementation manner of the first aspect of the present invention, use the CP decomposition based on ALS for tensor filling, alternately update the factor matrix, and continuously fit the original tensor.

[0012] Optionally, in the fifth implementation manner of the first aspect of the present invention, when updating one of the factor matrices, put the other two factors into the shared memory, and then extract the corresponding valid values required in the factor matrix from the shared memory.

[0013] Optionally, in the sixth implementation manner of the first aspect of the present invention, a corresponding matrix after a sub-tensor is matrixized is processed by one block, each row of the matrix is processed by the threads of this block, and after all blocks are calculated, the local results of all processing units are summarized to obtain the final reduction result.

[0014] The second aspect of the present invention provides a parallel tensor filling system based on ALS, and the system includes: A generation module, configured to collect feature information of a data set to be filled, and establish a tensor according to preset data features, and generate an initial factor matrix of the tensor; A division module, configured to obtain the number and positions of valid values of the tensor, divide the sub-tensors according to the performance of the hardware device, obtain a plurality of sub-tensors for parallel processing, and store the sub-tensors in the COO format; An update module, configured to allocate a processing unit to a sub-tensor, and perform a reduction operation, and assign the updated values of each row to the factor matrix being updated; An alternating optimization module, configured to alternately update three factor matrices, perform multiple rounds of iteration, after each round of iteration, perform an outer product on the updated factor matrices to obtain a filled tensor, calculate the loss between the front and back tensors, and alternately optimize each factor matrix to implement missing value filling.

[0015] The third aspect of the present invention provides a parallel tensor filling device based on ALS. The parallel tensor filling device based on ALS includes a memory and at least one processor, and instructions are stored in the memory; the at least one processor calls the instructions in the memory so that the parallel tensor filling device based on ALS executes each step of the parallel tensor filling method described in any one of the above.

[0016] The fourth aspect of the present invention provides a computer-readable storage medium, and instructions are stored on the computer-readable storage medium, and when the instructions are executed by a processor, each step of the parallel tensor filling method described in any one of the above is implemented.

[0017] Compared with the prior art, the present invention has at least the following advantages or beneficial effects: 1. The present invention proposes a parallel tensor completion method based on ALS. The steps include: constructing a tensor according to data characteristics and initializing factor matrices; extracting the valid values of the tensor and their positions; dividing the tensor into multiple sub-tensors according to the hardware performance, storing them in COO format and processing them in parallel; during the processing, placing the fixed factor matrices in the shared memory to accelerate access; after each sub-tensor is processed, reducing and updating the factor matrices; alternately updating the three factor matrices and performing multiple rounds of iteration until the error converges or reaches the iteration upper limit, and finally reconstructing the filled tensor through outer product. This method significantly improves the parallel computing efficiency, reduces the memory access overhead, and speeds up the convergence rate through reasonable data division and shared memory optimization, and is applicable to the efficient completion of large-scale sparse tensors; 2. By dividing the original tensor, the present invention makes the valid values of each factor matrix differ very little, so that the valid values to be processed by the computing units for subsequent sub-tensors differ little, realizing load balancing and avoiding load imbalance caused by directly dividing by rows. This helps to reduce the waiting time between different computing units, improve the parallel processing efficiency, reduce the overall computing latency, and further enhance the scalability and stability of the system in the scenario of large-scale sparse tensor processing; 3. The present invention shortens the access time during calculation by using shared memory to store factor matrices. Since the access speed to shared memory is much higher than that of global memory, the performance can be improved by about a hundred times. During the process of updating the factor matrices, each sub-tensor processed in parallel needs to frequently access the corresponding rows of the other two factor matrices when updating each row of the current factor matrix. After loading these two fixed factor matrices into the shared memory, the number of global memory accesses can be significantly reduced, the memory bandwidth pressure can be reduced, and thus the iteration speed can be accelerated. This optimization strategy not only improves the data access efficiency, but also enhances the execution performance of the overall algorithm in the GPU parallel architecture, especially applicable to scenarios with high-frequency sparse access. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.

[0019] Figure 1 It is a flowchart of the parallel tensor completion method based on ALS provided by the embodiment of the present invention; Figure 2 It is a schematic diagram of ALS updating factor matrices provided by the embodiment of the present invention; Figure 3 It is a schematic diagram of uniformly dividing the original tensor provided by the embodiment of the present invention; Figure 4 It is a schematic diagram of the GPU memory architecture provided by the embodiment of the present invention; Figure 5Schematic diagram of dividing the original tensor based on thermal conduction for the embodiments of the present invention; Figure 6 Schematic diagram of COO compressed storage for the embodiments of the present invention; Figure 7 Schematic diagram of the parallel acceleration mechanism based on GPU shared memory for the embodiments of the present invention; Figure 8 MAE comparison of different datasets under different algorithms at different sampling rates for the embodiments of the present invention; Figure 9 RMSE comparison of different datasets under different algorithms at different sampling rates for the embodiments of the present invention; Figure 10 Memory comparison of different datasets under different algorithms at different sampling rates for the embodiments of the present invention; Figure 11 Time comparison of different datasets under different algorithms at different sampling rates for the embodiments of the present invention; Figure 12 Schematic diagram of the structure of the parallel tensor filling system based on ALS for the embodiments of the present invention; Figure 13 Schematic diagram of the structure of the parallel tensor filling device based on ALS for the embodiments of the present invention. Detailed implementation manners

[0020] The technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. It should be understood that the embodiments are only used to illustrate a part of the specific implementation manners of the present invention, rather than to limit the protection scope of the present invention. For those of ordinary skill in the art, various deformations, substitutions or equivalent improvements made without departing from the essence of the present invention shall fall within the protection scope of the present invention.

[0021] For ease of understanding, the specific process of the embodiments of the present invention will be described below. Please refer to Figure 1 The flowchart of parallel tensor filling based on ALS provided by the embodiments of the present invention. The method specifically includes the following steps: Step 101: Collect the feature information of the dataset to be filled, and establish a tensor according to the preset data features, and generate the initial factor matrix of the tensor; Step 102: Obtain the number and positions of the valid values of the tensor, divide the sub-tensors according to the performance of the hardware device, obtain multiple sub-tensors for parallel processing, and store the sub-tensors in the COO format; Step 103: Allocate a processing unit for a sub-tensor, and perform a reduction operation, and assign the updated value of each row to the factor matrix being updated currently; Step 104: Alternately update the three factor matrices and perform multiple rounds of iteration. After each round of iteration, calculate the outer product of the updated factor matrices to obtain the filled tensor, calculate the loss between the tensors before and after, and alternately optimize each factor matrix to achieve missing value filling.

[0022] Based on the low-rank assumption of multi-dimensional data, the tensor completion problem usually relies on the core tool CANDECOMP / PARAFAC (CP) decomposition. CP decomposition represents the tensor as the sum of a finite number of rank-one tensors: ; In the formula , and correspond to the factor matrices in modes 1, 2, and 3 respectively. , , are the vectors of the r-th rank-one component of the tensor in the three modes respectively. In the ALS framework, the factor matrices are updated mode by mode to minimize the reconstruction error. The key significance of CP decomposition lies in that through low-rank approximation (i.e., ), the potential multi-dimensional structure of the tensor can be compressed and represented as a combination of several factor matrices, thus providing a parameterized basis for subsequent missing value recovery.

[0023] Based on the above CP decomposition model, the tensor completion problem can be formulated as a constrained optimization task. Given a partially observed tensor (the observed positions are marked by the index set ), the goal is to find a complete tensor such that it is consistent with the observed values at , and satisfies the low-rank constraint: ; In the formula, represents the projection operator (retaining the elements in and setting other elements to zero), represents the CP rank, is a predefined rank parameter. This optimization problem fundamentally uses the low-rank prior to reconstruct the missing elements from the partial observations. CP decomposition parameterizes the optimization variable as the joint action of factor matrices, transforming the tensor completion problem into the , and update problem.

[0024] ALS solves this optimization problem iteratively by alternately fixing some variables and updating the others: First, update the factor matrix , as shown in the following formula, fix and , and expand into a matrix according to Pattern 1 : ; where represents and 's Khatri-Rao product, represents the pseudo-inverse; Then fix the remaining factor matrices respectively, expand and update them according to Pattern 2 and Pattern 3 respectively and , and the specific formulas are as follows: ; ; Repeat the above process until convergence. The advantage of ALS is that the updates of each modality are independent and easy to be parallel accelerated on the GPU.

[0025] Refer to Figure 2 shown. When updating the x-th row (x, :) of the factor matrix , perform a mode-1 slice on the original tensor, and use the (x, :, :) forward slice of the original tensor all valid values in , for calculation; when updating the y-th row (y, :) of the factor matrix , perform a mode-2 slice on the original tensor, and use the (:, y, :) vertical slice of the original tensor matrix all valid values in , for calculation; similarly, when updating the z-th row (z, :) of the factor matrix , perform a mode-3 slice on the original tensor, and use the (:, :, z) forward slice of the original tensor matrix all valid values in , for calculation. Since the update of each factor matrix only depends on the valid values of the specific mode slice, these update operations can be independently and parallelly executed. This feature enables the use of the advantage of parallel computing to accelerate the update of the factor matrix when filling tensors with ALS.

[0026] Due to the dynamic change of the network topology and the sampling difference of the monitoring devices, the network traffic data naturally exhibits strong sparsity and non-uniform distribution characteristics. As Figure 3 shown, the original tensor The valid values (such as traffic peaks and abnormal events) are often concentrated in a few local subtensors (such as specific time periods or routing nodes), while most regions are missing values or low values. This distribution characteristic poses a severe challenge to parallel tensor filling based on ALS.

[0027] Traditional GPU parallelization methods adopt an equal division strategy, evenly splitting the tensor into subtensors of the same size (such as horizontal slices where each subtensor is (1, J, K) in Mode 1), and independently processed by different GPU thread blocks. However, due to the highly uneven distribution of valid values, some subtensors may contain dense valid values (such as a subtensor containing 452 observation points), while the computational load of other subtensors (such as only containing 238 observation points) is significantly reduced. This load imbalance leads to two key problems: 1. Thread block idle waiting: GPU thread blocks execute in a synchronous manner and cannot enter the next iteration until all subtensors are processed. As Figure 2 shown, the thread block (Thread Block 2) processing completes the calculation quickly due to the light load, but needs to wait for the thread block (Thread Block 1) processing to run for a long time, resulting in waste of hardware resources; 2. Iterative convergence delay: The overall convergence speed of the ALS algorithm is limited by the execution time of the slowest thread block. When the proportion of sparse subtensors is high, frequent thread waiting significantly increases the time consumption of a single iteration, especially when dealing with large-scale tensors (such as I = ), and the time cost increases exponentially.

[0028] Existing methods (such as cuTC and TensorD) compress the data scale through sparse storage, but their static division strategy does not consider the spatial locality of valid values, resulting in the decoupling of computational resource utilization and data distribution. For example, in the scenario of bursty traffic monitoring, traditional division may split the instantaneous peak region into multiple subtensors, forcing a single thread block to be unable to completely capture local features, further exacerbating load imbalance and convergence instability. Therefore, designing a dynamic and adaptive subtensor division method to achieve load balance among thread blocks has become the core issue in improving ALS parallel efficiency.

[0029] To obtain results quickly, parallel computing is undoubtedly a very effective option. With the progress of technology, especially in the fields of data processing and machine learning, the application of GPUs (Graphics Processing Units) has become increasingly widespread. This is because GPUs have powerful parallel processing capabilities and can execute a large number of computing tasks simultaneously, thus significantly improving the processing speed. Compared with traditional CPUs, GPUs are more efficient in performing complex calculations, especially in scenarios that require processing large amounts of data. For example, in the fields of image processing, scientific computing, and deep learning, the parallel architecture of GPUs enables multiple computing units to work simultaneously, greatly reducing the computing time. Therefore, using GPUs for parallel computing has become a common and effective solution in modern computing tasks.

[0030] The Figure 4 following shows the approximate thread structure and memory structure of a GPU. A GPU is divided into multiple Blocks, and each Block contains multiple Threads. A Thread is the smallest execution unit of a GPU. A Thread has its own local memory and register memory. Register memory is the smallest memory in a GPU, but it responds immediately. Each Block has a shared memory, which is a memory area that all threads in the Block can access. It is also small, but its response speed is only 4 - 6 clocks. Global memory is the largest memory in a GPU. Generally, uploading data to a GPU means uploading it to global memory, but its access time is 100 times slower than accessing shared memory.

[0031] Although the parallel computing ability of GPUs provides a hardware acceleration foundation for tensor filling, the deficiencies in memory access patterns and data locality optimization of existing methods severely limit the actual performance, specifically manifested as the following two aspects of problems: 1. Risk of intermediate data video memory explosion: During the ALS iteration process, it is necessary to calculate the Khatri - Rao product, and its explicit storage places extremely high requirements on the GPU video memory capacity. For example, when J = 15,000, K = 29,000, and R = 100, this matrix occupies approximately 16GB of video memory (in float32 format), and a single iteration may involve multiple such intermediate matrices. For devices with limited video memory (such as the RTX4090 which only has 24GB in total), this problem forces the algorithm to reduce the rank R or use low - precision calculations, directly damaging the filling accuracy; 2. Insufficient utilization of data locality: The fast access feature of GPU shared memory is not effectively utilized. For example, when threads within the same Block process adjacent data, they need to repeatedly access specific row data of B and C, but existing methods do not cache these frequently accessed data in the shared memory, resulting in redundant global memory accesses. Theoretical analysis shows that if the local rows of the factor matrix are cached in the shared memory, the data access latency can be reduced to 1% of the original.

[0032] Existing optimization schemes (such as GIGATENSOR) avoid intermediate matrix storage through implicit calculations, but their element-by-element calculation mode does not combine with the characteristics of the GPU memory hierarchy, resulting in a significant increase in thread synchronization overhead. In addition, although sparse storage formats (such as CSR, COO) reduce video memory occupancy, they require additional parsing of index data, further exacerbating the computational complexity. Therefore, designing a parallel acceleration mechanism that takes into account both video memory efficiency and data locality has become a key challenge in improving GPU utilization.

[0033] To address the above challenges, this study proposes the following core optimization schemes: Thermal conduction load balancing: Based on the spatial distribution characteristics of the effective values, an adaptive sub-tensor partitioning algorithm is designed to dynamically adjust the computational load of each GPU thread block, ensuring an even distribution of thread tasks between sparse and dense regions; Shared memory mechanism: Cache the factor matrix data that is frequently accessed through shared memory to reduce the global memory access latency. At the same time, use an implicit element-by-element calculation mode to replace the explicit intermediate matrix storage, effectively alleviating the video memory pressure.

[0034] Next, the above two methods will be introduced in further detail: Please refer to Figure 5 Based on the schematic diagram of partitioning the original tensor by thermal conduction, this paper proposes a thermal conduction-based load balancing method. Its core idea draws on the physical process of heat diffusion from high-density regions to low-density regions in heat conduction, and realizes the load migration of computational tasks by dynamically adjusting the sub-tensor boundaries. Specifically: Load characterization: Use the number of effective values of the sub-tensor as the computational load indicator; Load diffusion mechanism: Expand the adjacent slices of the high-load sub-tensor into the low-load region, and redistribute the effective values by adjusting the slice thickness (such as expanding along the j-axis in the vertical slice mode), so that the load of the adjusted sub-tensor satisfies ; is the tolerance parameter. Keep the original spatial positions of the effective values unchanged and only change the sub-tensor partitioning boundaries to ensure that the row indices of the factor matrix during ALS update strictly correspond to the original tensor, avoiding computational logic errors.

[0035] Boundary constraint: During the adjustment process, keep the original spatial positions of the effective values unchanged and only change the sub-tensor partitioning boundaries to ensure that the row indices of the factor matrix during ALS update strictly correspond to the original tensor, avoiding computational logic errors.

[0036] Through the above dynamic adjustment, the computing loads of each GPU thread block tend to be balanced, maximizing the utilization of parallel computing resources, thereby reducing the thread idle waiting time and improving the overall iteration efficiency.

[0037] To make the parallelism more sufficient, as many sub-tensors as possible should be set. However, the number of valid values in the sub-tensors cannot be too small. If the sub-tensors all become slices of the smallest unit in that dimension, the partitioning will not make much sense. The size is related to the valid values of the original tensor and the hardware device. When mode is 1, it means slicing the tensor horizontally, and the diffusion is ; when mode is 2, it slices the tensor vertically, and the diffusion is ; when mode is 3, it slices the tensor frontally, that is ; ; In the for loop, the sub-tensors are partitioned in sequence, adding one slice at a time. There is a nested while loop inside the for loop. Here, it is calculated whether the number of valid values in the slice is closest to the target number of valid values as much as possible. Generally, it is almost impossible for the number of valid values to be exactly equal to the target number of valid values because the smallest unit for expansion or contraction is a sub-tensor of (1, J, K), and the valid values inside are uncertain. Similarly, load migration can be performed on vertical slices and frontal slices until the loads of each sub-tensor are balanced within a certain range.

[0038] According to the thermal load balancing method, the non-zero values are partitioned to obtain one or more batches of sub-tensors. These sub-tensors update the corresponding rows respectively and can be processed in parallel. Each GPU block corresponds to a sub-tensor, and multiple blocks can be processed simultaneously. Of course, a mask needs to be added during calculation to distinguish between valid values and zero values.

[0039] Please refer to Figure 6 , for sparse matrices, first, COO compression storage can be adopted. The COO format is a common storage method for sparse tensors or sparse matrices, mainly used to save the positions of non-zero elements and their corresponding values. In the COO format, the data is stored through three corresponding arrays: the row index array (row indices), the column index array (column indices), and the value array (values). For multi-dimensional tensors, it will be extended to multiple coordinate arrays. For example, for a three-dimensional tensor, i_indices, j_indices, and k_indices are used to represent the indices in the three dimensions respectively. Specifically, as Figure 5As shown, the sub-tensor is first matrixized and unfolded according to mode-1, mode-2, and mode-3 respectively. The position of each non-zero element is jointly determined by its coordinates in each dimension, and its corresponding value is stored in the values array. Therefore, the information of the nth non-zero element can be completely described by row[n], col[n], and values[n] (or multiple coordinate arrays of high-dimensional tensors). This structure is concise and intuitive, facilitating the traversal and reconstruction of the original data.

[0040] According to the definition of the Kronecker product, The dimension of is and When they are relatively large, its dimension may reach hundreds of thousands or even millions, thus bringing huge storage overhead. It is very difficult to directly store and calculate such a large amount of intermediate results on the GPU. GigaTensor can be used to avoid the problem of data explosion. It is mentioned in the GigaTensor paper that updating the factor matrix A can be divided into three steps: ; ; ; To solve the problem of the explosion of intermediate data caused by the explicit calculation of the matrix in Step1, GigaTensor no longer explicitly constructs , but only constructs , , , and then performs element-by-element calculation: use a for loop to divide into , , and calculate them in three steps, which greatly reduces the memory space overhead; ; ; ; After dividing the valid values according to the heat conduction load balancing method, a group of sub-tensors is obtained. It is known that these sub-tensors update the corresponding rows respectively, so these sub-tensors can be processed in parallel, that is, one block in the GPU corresponds to processing one sub-tensor, and multiple blocks are processed in parallel.

[0041] Such as Figure 7As shown, the main calculation to be optimized is that of M1. According to the formula, three for loops are required to update the matrix M1. If the first-level loop selects r, it means updating by column. It can be found that once r is determined (as long as it is within the range), it will not affect the calculations of the second and third-level loops. Therefore, the first-level loop can be opened, allowing one thread in the block to process one r value, and all threads can also process in parallel, further leveraging the parallel characteristics of the GPU to reduce loop nesting. In addition, the factor matrices B and C are used multiple times in the innermost loop for calculations, so the factor matrices B and C can be placed in the shared memory of the block. Of course, the results processed by each thread for one column need to be stored in an array, so this array can also be placed in the shared memory for easy access by each thread.

[0042] The key lies in the writing of the kernel function. The specific idea is as follows: First, initialize bx and tx as the current block number and the current thread number respectively; then copy the factor matrices B, C, and the array for storing the results of the current block from the global memory to the shared memory, and the tx-thread processes the tx-th column. Since the number of columns of the matrix is fixed as R, the number of threads required cannot exceed R. In the first-level for loop, the elements of all rows in this column are processed in sequence; in the second-level for loop, the accumulation in the corresponding formula is performed, and specific elements in the i-th row of the factor matrices B and C (the -th column of B and the -th column of C) are required for calculation, which is specifically divided into three steps 、 、 for calculation. is to accumulate the results of multiplying and , and add as the regularization parameter to prevent overfitting; after each thread completes the processing of the i-th row, it will obtain a , which is placed in a specific position in the array. Since each thread processes a different column, and completing one second-level for loop only processes a specific row, there will be no conflict among all threads.

[0043] It should be explained that the above filling is to fill the missing data in the dataset. The specific process is as follows: 1) Repeated iteration: After updating the three factor matrices each time, perform an outer product on the three factor matrices to calculate the filled tensor, and calculate the relative root mean square error (RMSE) between the filled tensor and the original tensor: ; and the relative mean absolute error (MAE): ; In the formula, , , and . represents the total number of elements, represents the , , -th data in the original tensor, represents the , , -th data in the predicted tensor. is the index set of missing values, is the total number of missing values. Repeat this process. There are usually two stopping conditions: one is that the difference in the recovery loss (such as mean square error and other measurement metrics) between two consecutive iterations is less than a given threshold, which is set to in the embodiments of the present invention; the other is reaching the maximum number of iterations. When any of these stopping conditions is met, the iteration terminates.

[0044] 2) Obtain the converged factor matrices When the iteration process above converges, the final factor matrices A, B, and C are obtained.

[0045] 3) Recover the tensor using the factor matrices According to the formula , substitute the converged factor matrices A, B, and C to calculate the values of the missing terms in the original tensor. Here, i, j, and k traverse the index range of the entire tensor. For the positions of the originally missing data, their recovered values are obtained through calculation, thereby achieving the complete recovery of the entire original tensor data and obtaining the complete and filled tensor data, which can be used for subsequent analysis and applications.

[0046] As Figures 8 - 11 shown, the experimental hardware is NVIDIA GeForce RTX 4090 GPU. In a specific implementation, six publicly available real datasets are used to verify the effectiveness of the algorithms involved in the embodiments of the present invention, including the Chengdu traffic dataset with a size of 2016*128*64, the Beijing traffic dataset with a size of 1530*128*128, the New York taxi dataset NYC with a size of 30*30*1464, the traffic dataset PeMS with a size of 228*288*44, the ECW dataset 08 with a size of 797*720*16, and the dataset 09 with a size of 1022*720*16. The six datasets are unevenly sampled to simulate the experimental conditions of data imbalance, and the sampling rate ranges from 10% to 90%.

[0047] Evaluate the fast and efficient data recovery of GIGA_ALS from six aspects: the precision convergence of the balanced data volume partitioning method for GIGA_ALS at different sampling rates and the impact of the high-dimensional structure of GIGA_ALS on precision. By comparing GIGA_ALS with TensorD_ALS (CP decomposition using TensorFlow), cuTC_sgd, cuTC_ALS, and pyCP_APR (alternating Poisson regression), all of these four methods utilize the parallel characteristics of the GPU for CP decomposition.

[0048] To verify the effectiveness of the load balancing strategy, compare the time consumption and speedup ratio between not using the strategy (Time1) and after using it (Time2). ; When not using thermal conduction load balancing, one block processes a sub-tensor of a minimum unit slice, and the time is Time1; after using thermal conduction load balancing, one block processes a sub-tensor (composed of multiple minimum unit slices), and the time is Time2. At this time, shared memory is not used to evaluate the contribution of thermal conduction load balancing to the model performance.

[0049] The speedup ratios of dataset ECW_08 using thermal conduction load balancing at different sampling rates are 2.002, 2.553, 2.882, 3.033, 3.136, 3.457, 3.671, 3.661, 3.865 respectively; the speedup ratios of dataset ECW_09 using thermal conduction load balancing at different sampling rates are 2.487, 2.938, 3.404, 3.687, 3.815, 3.843, 4.146, 4.146, 3.290 respectively; the speedup ratios of dataset CDTaxi using thermal conduction load balancing at different sampling rates are 1.55, 1.82, 2.05, 2.18, 2.39, 1.69, 1.92, 1.51, 1.98 respectively; the speedup ratios of dataset NYC using thermal conduction load balancing at different sampling rates are 1.428, 1.136, 1.384, 1.375, 1.382, 1.356, 1.281, 1.333, 1.315 respectively; the speedup ratios of dataset PeMS using thermal conduction load balancing at different sampling rates are 1.182, 1.202, 1.352, 1.545, 1.545, 1.630, 1.520, 1.486, 1.576 respectively; the speedup ratios of dataset BJTaxi using thermal conduction load balancing at different sampling rates are 1.389, 1.724, 1.753, 1.718, 1.511, 1.692, 1.829, 2.412, 2.139 respectively.

[0050] Key conclusions: High-sparsity scenarios have significant advantages: The load balancing strategy enables speedup ratios of up to 3.865 and 4.146, indicating that it effectively reduces the load imbalance caused by "long strips"; Generalizability verification: For dense datasets (such as NYC and PeMS), the strategy can still improve efficiency (speedup ratios of 1.315 - 4.052), but the improvement amplitude is affected by data distribution. For example, the speedup ratio of NYC is 1.281 at a sampling rate of 0.7, while that of PeMS reaches 4.052 at 0.9, indicating that the strategy is sensitive to data locality; Influence of load balancing threshold: When the valid value of the sub-tensor approaches the set threshold (ε), the utilization rate of computing resources is optimal. For example, the speedup ratio of CDTaxi drops suddenly (1.69 → 1.125) at a sampling rate of 0.5, reflecting that the threshold needs to be adjusted dynamically according to the data (the thresholds on each dimension of the same dataset are fixed in the experiment).

[0051] The speedup ratios of dataset ECW_08 using the shared memory acceleration mechanism at different sampling rates are 1.375, 1.255, 1.472, 1.485, 1.566, 1.677, 1.465, 1.580, 1.528 respectively; The speedup ratios of dataset ECW_09 using the shared memory acceleration mechanism at different sampling rates are 1.047, 1.262, 1.409, 1.432, 1.380, 1.522, 1.603, 1.517, 1.445 respectively; The speedup ratios of dataset CDTaxi using the shared memory acceleration mechanism at different sampling rates are 1.101, 1.250, 1.370, 1.339, 1.192, 1.125, 1.176, 1.415, 1.427 respectively; The speedup ratios of dataset NYC using the shared memory acceleration mechanism at different sampling rates are 1.838, 1.861, 1.749, 1.774, 1.764, 1.826, 2.058, 2.064, 1.987 respectively; The speedup ratios of dataset PeMS using the shared memory acceleration mechanism at different sampling rates are 1.713, 2.067, 1.860, 1.961, 2.195, 3.022, 3.747, 4.070, 4.052 respectively; The speedup ratios of dataset BJTaxi using the shared memory acceleration mechanism at different sampling rates are 1.090, 1.078, 1.143, 1.287, 1.511, 1.277, 1.280, 1.088, 1.225 respectively; Key conclusions: Significantly reduce the memory bandwidth pressure: The shared memory mechanism enables the PeMS dataset to achieve an acceleration ratio of 4.052 at a sampling rate of 0.9, indicating a significant improvement in the memory access efficiency of high-frequency data; Limited benefits for sparse data: The acceleration ratio of BJTaxi is only 1.090 at a sampling rate of 0.1. Due to the insufficient utilization of shared memory under sparse data, it is necessary to combine load balancing strategies for collaborative optimization; Memory-computation balance: The acceleration ratios of ECW_08 and NYC increase with the increase of the sampling rate (1.375 → 1.528, 1.838 → 1.987), indicating that shared memory has more advantages in dense computing scenarios.

[0052] It can be seen from Figures 7 - 8 that the advantages of MAE and RMSE: GIGA_ALS has the lowest error on datasets such as ECW_08 and PeMS (MAE is reduced by 10% - 15%) because the heat conduction strategy reduces the local overfitting caused by load imbalance; Robustness verification: In CDTaxi and NYC, the errors of each algorithm are similar, but GIGA_ALS has the fastest convergence speed (the number of iterations is reduced by 20%), thanks to the shared memory for accelerating matrix updates.

[0053] Under the condition that the loss is basically the same, the memory and time are compared again ( Figures 9 - 10 ). TensorD does not extract the valid values and their positions during calculation, but replaces the missing values with 0, and 0 also occupies memory, so the GPU memory of TensorD will not change with the sampling rate. Because cuTC preprocesses the data in advance and uses sparse matrix storage, the memory occupation is relatively small, while SGD processes only one or a few samples at a time instead of the entire dataset, which means that the required GPU memory is the least when calculating gradients and updating model parameters. Results of efficiency and resource comparison: Memory occupation: Taking the CDTaxi dataset as an example, the memory occupation of GIGA_ALS increases from 212MB to 308MB, an increase of 96MB, which is much smaller than other algorithms. This stable growth indicates that the algorithm has stronger adaptability to sampling rate changes and can maintain a lower memory occupation at different sampling rates. Compared with TensorD, the average memory occupation is reduced by 80.45%, and the maximum memory occupation is reduced by 84.06%; Time performance: At different sampling rates of the four datasets, the time acceleration ratio of this algorithm and the other three benchmark algorithms is as high as 90%. The algorithm in this paper significantly improves the computing efficiency while ensuring the accuracy, and is suitable for large-scale network traffic monitoring scenarios with high real-time requirements and uneven data distribution.

[0054] Please refer to Figure 12 , the structural schematic diagram of the parallel tensor filling system based on ALS provided by the embodiment of the present invention. The system includes: A generation module, configured to collect feature information of a dataset to be filled, establish a tensor according to preset data features, and generate an initial factor matrix of the tensor; A partitioning module, configured to obtain the number and positions of valid values of a tensor, partition the tensor into sub-tensors according to the performance of a hardware device, obtain multiple sub-tensors for parallel processing, and store the sub-tensors in a COO format; An update module, configured to allocate a processing unit to a sub-tensor and perform a reduction operation, and assign the updated values of each row to the factor matrix being updated currently; An alternating optimization module, configured to alternately update three factor matrices through multiple rounds of iteration. After each round of iteration, an outer product of the updated factor matrices is calculated to obtain a filled tensor, the loss between the front and back tensors is calculated, and each factor matrix is alternately optimized to implement missing value filling.

[0055] Figure 13 FIG. 10 is a schematic structural diagram of a parallel tensor filling device based on ALS provided by an embodiment of the present invention. The parallel tensor filling device 600 based on ALS may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) 610 (for example, one or more processors) and a memory 620, and one or more storage media 630 for storing application programs 633 or data 632 (for example, one or more mass storage devices). Among them, the memory 620 and the storage media 630 may be transient storage or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the parallel tensor filling device 600 based on ALS. Further, the processor 610 may be configured to communicate with the storage media 630 and execute a series of instruction operations in the storage media 630 on the parallel tensor filling device 600 to implement the method provided in the above embodiment.

[0056] The parallel tensor filling device 600 based on ALS may further include one or more power supplies 640, one or more wired or wireless network interfaces 650, one or more input / output interfaces 660, and / or one or more operating devices 631, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 13 the shown structure of the parallel tensor filling device based on ALS does not limit the computer device provided by the present invention, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0057] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the respective steps of the ALS-based parallel tensor filling method provided in the above-described embodiments.

[0058] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices, apparatuses, or units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0059] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0060] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A parallel tensor completion method based on ALS, characterized in that The method includes the following steps: Collect the feature information of the dataset to be filled, and establish a tensor according to the preset data features, and generate the initial factor matrix of the tensor; Obtain the number and position of the valid values of the tensor, divide the sub-tensors according to the performance of the hardware device, obtain multiple sub-tensors for parallel processing, and store the sub-tensors in the COO format; Allocate a processing unit for a sub-tensor, and perform a reduction operation, and assign the updated value of each row to the currently updated factor matrix; Alternately update the three factor matrices, perform multiple rounds of iteration. After each round of iteration, perform an outer product on the updated factor matrices to obtain the filled tensor, calculate the loss between the front and back tensors, and alternately optimize each factor matrix to achieve missing value filling.

2. The parallel tensor filling method based on ALS according to claim 1, wherein, Obtain three features in the dataset to be filled that have real physical relationships with each other, establish a corresponding three-dimensional tensor, fill the missing values with 0 values, and perform feature scaling using the maximum-minimum normalization to obtain the normalized tensor.

3. The parallel tensor completion method based on ALS according to claim 1, wherein Realize the load migration of the computing task by dynamically adjusting the boundaries of the sub-tensors. Use the valid values of the sub-tensors as the computing load index, expand the adjacent slices of the high-load sub-tensors to the low-load area, and re-allocate the valid values by adjusting the slice thickness, so that the load of the adjusted sub-tensors meets: ; Among them, is the tolerance parameter, is the number of optimal valid values for each subtensor, is the payload value of the subtensor.

4. The parallel tensor filling method based on ALS according to claim 1, wherein Traverse the entire factor matrix to find all non-zero elements. For each non-zero element, record its row index, column index, and corresponding value, and store these indexes and values in three arrays respectively.

5. A parallel tensor filling method based on ALS according to claim 1, characterized in that, Use CP decomposition based on ALS for tensor filling, alternately update the factor matrices, and continuously fit the original tensor.

6. The parallel tensor filling method based on ALS according to claim 5, characterized in that, When updating one of the factor matrices, place the other two factor matrices in the shared memory, and then extract the required valid values from the factor matrices in the shared memory.

7. A parallel tensor filling method based on ALS according to claim 1, characterized in that The corresponding matrix after matrixizing a sub-tensor matrix is processed by a block. Each row of the matrix is processed by the threads of this block. After all blocks are calculated, the local results of all processing units are summarized to obtain the final reduction result.

8. A parallel tensor completion system based on ALS, characterized in that, The system includes: A generation module, which is used to collect the feature information of the dataset to be filled, and establish a tensor according to the preset data features, and generate the initial factor matrix of the tensor; A division module, which is used to obtain the number and position of the valid values of the tensor, divide the sub-tensors according to the performance of the hardware device, obtain multiple sub-tensors for parallel processing, and store the sub-tensors in the COO format; An update module, which is used to allocate a processing unit for a sub-tensor, and perform a reduction operation, and assign the updated value of each row to the currently updated factor matrix; An alternating optimization module, which is used to alternately update the three factor matrices, perform multiple rounds of iteration. After each round of iteration, perform an outer product on the updated factor matrices to obtain the filled tensor, calculate the loss between the front and back tensors, and alternately optimize each factor matrix to achieve missing value filling.

9. A parallel tensor filling device based on ALS, characterized in that, The parallel tensor filling device based on ALS includes a memory and at least one processor, and instructions are stored in the memory; the at least one processor calls the instructions in the memory to cause the parallel tensor filling device based on ALS to execute each step of the parallel tensor filling method according to any one of claims 1-7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed by the processor, each step of the parallel tensor filling method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Multi-tensor visual data filling method based on convex optimization

    CN108765517A

  • A tensor decomposition and reconstruction method based on GPU

    CN109033030A

  • Network traffic data filling method, device and equipment and storage medium

    CN110941793A

  • Sparse tensor canonical decomposition method based on data division and calculation distribution

    CN112765094A

  • Sparse tensor canonical decomposition method and system based on adaptive load distribution

    CN117311971A

Cited By

  • Data read-write method, processor, electronic equipment and storage medium

    CN120762918A

  • Data reading and writing method, processor, electronic device, and storage medium

    CN120762918B

  • Method, computing device, medium and program product for performing reduction computation

    CN120849770A

  • Sparse tensor parallel filling method and system based on task division and dynamic scheduling

    CN122173305A

  • A Parallel Sparse Tensor Filling Method and System Based on Task Partitioning and Dynamic Scheduling

    CN122173305B