An incremental iterative approach for distributed deep learning
By constructing DAG graphs in a heterogeneous distributed environment, filtering RDD data sets and performing iterative calculations on the GPU, the resource utilization problem of deep learning incremental iteration models when processing streaming data in a heterogeneous environment is solved, and efficient data processing is achieved.
Patent Information
- Application Number
- CN202211210718.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In a heterogeneous distributed environment with integrated GPUs, building a deep learning incremental iterative model faces challenges, especially how to rationally utilize time and resources when processing streaming data.
By constructing a DAG graph, filtering the RDD data set, converting it into a data type that the GPU can process, and iteratively calculates in the GPU global video memory, using shared memory for incremental iterative calculations, and optimizing the utilization of memory and computing resources.
It effectively reduces repeated calculations, reduces the total calculation time, and improves the efficiency of deep learning data processing.
Smart Images

Figure CN115509512B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to an incremental iterative method for distributed deep learning. Background Art
[0002] With the resurgence of artificial intelligence enthusiasm, parallel processing platforms dedicated to deep learning have become the focus of many researchers. As mainstream representatives of the MapReduce programming model, Flink and Spark are very suitable for computationally intensive data analysis and iterative computing applications. However, there are still many different characteristics between distributed computing frameworks and GPUs, which brings challenges to building deep learning incremental iterative models in heterogeneous distributed environments with integrated GPUs.
[0003] At the same time, in actual application scenarios, streaming data changes dynamically in real time. How to reasonably utilize time and resources for the calculation of incremental data generated when the data changes is an issue that traditional big data processing methods urgently need to address. Summary of the invention
[0004] The purpose of the present invention is to provide an incremental iterative method for distributed deep learning, which can effectively complete the task of efficient data processing.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] An incremental iterative method for distributed deep learning, comprising:
[0007] Step 1: Construct DAG. Get the RDD ID and the dependencies between RDDs, and construct a directed acyclic graph consisting of multiple triplets including head RDD, dependency, and tail RDD.
[0008] Step 2: Filter the RDD data set. In order to increase the subsequent GPU utilization efficiency, when the memory reaches a certain threshold, the weights of all RDD data in step 1 are calculated and filtered to filter out RDDs with small weights. The filtered data set is the RDD data set that needs to be cached in memory.
[0009] Step 3: Convert the RDD data set filtered in step 2 above into a data type that can be processed by the GPU and store it in the GPU global memory.
[0010] Step 4: Perform iterative calculation. Allocate a data set to each thread block and place it in the shared memory. The data set is a collection of data with high local access frequency in each iteration. Each thread in the thread block reads the data to be calculated in the global video memory in turn and performs iterative calculation. The data set stored in the shared memory is updated with the iteratively calculated data set.
[0011] Step 5: Perform incremental iterative calculations. When the data set undergoes incremental changes, the shared memory reads the incremental data from the global video memory and performs incremental iterative calculations to obtain the incrementally iterated data set, which is then used to update the data set in the shared memory.
[0012] Furthermore, in step 1, the specific method of constructing DAG includes:
[0013] Step 1.1: Input sample data, traverse all RDD function operations, and obtain the triple t consisting of all input RDDs, dependencies, and output RDDs:
[0014] t=R RhR ⊕r⊕R Rt
[0015] Where R RhR is the input RDD ID, r is the dependency, R RtR is the output RDD ID. At the same time, the RDD that cannot form a triple relationship is discarded.
[0016] Step 1.2, read all triples t and construct DAG;
[0017] Furthermore, in step 2, the specific method of filtering the RDD data set includes:
[0018] Step 2.1: Determine whether the current memory storage capacity reaches the threshold. If so, continue with the following steps;
[0019] Step 2.2: Calculate the RDD weight. The formula for this process is defined as follows:
[0020]
[0021] Among them, w represents the weight value of the RDD partition, Indicates the computational cost of the RDD. Indicates the number of times the RDD is used, S p Indicates the size of the partition. Indicates the life cycle of the RDD. Indicates the position of the RDD input RDD for calculation. The elements in A = {α0, α1, α2, α3, α4} are constants, respectively and The normalized weight of the task is determined by the user's specific task requirements.
[0022] Furthermore, in step 4, the specific method for performing iterative calculation includes:
[0023] Step 4.1: Allocate a data set for each thread block in the global video memory and place it in the shared memory.
[0024] Step 4.2, each thread in the thread block reads the data to be calculated in the global video memory in turn, and performs iterative calculation;
[0025] Step 4.3, update the data set stored in the shared memory with the data set iteratively calculated in step 4.2;
[0026] Step 4.4: When the shared memory reaches a certain threshold, the data with low local access frequency in each iteration is transferred to the global video memory;
[0027] Furthermore, in step 5, the specific method of performing incremental iterative calculation includes:
[0028] Step 5.1: When the data set undergoes incremental changes, that is, when new iterative data is generated due to dynamic changes in the stream data, the shared memory reads the incremental data from the global video memory;
[0029] Step 5.2: perform incremental iterative calculation to obtain a data set after incremental iterative calculation. Incremental iteration is an iterative method for obtaining a new iterative result based on incremental data and original iterative results, and the incremental data is new iterative data generated due to business growth.
[0030] Step 5.3: Update the data set in the shared memory with the data set calculated by the incremental iteration in step 5.2 above. The data judged to be converged is stored in the global memory and is used in subsequent iterations without being updated, so as to achieve the purpose of incremental control.
[0031] Step 5.4: When the shared memory reaches a certain threshold, the infrequently accessed data is moved to the global video memory.
[0032] Advantages of the present invention:
[0033] The incremental iterative method of distributed deep learning mentioned above builds a DAG graph, filters and replaces the data in the memory according to the RDD weight, stores the data set in the global video memory of the GPU, and stores the data in the iterative calculation in the shared memory for calculation. It makes full use of the cache resources in the GPU and effectively reduces repeated calculations through calls between different storage structures, thereby reducing the total calculation time and improving the data processing efficiency of deep learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A flowchart of an incremental iterative method for distributed deep learning of the present invention;
[0035] Figure 2This is a structural diagram of the incremental iterative model proposed in the present invention. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical scheme and technical effect of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0037] like Figure 1 As shown, an incremental iterative method for distributed deep learning includes the following steps:
[0038] Step 1: Build DAG, specifically:
[0039] In step 1.1, this embodiment: input sample data, traverse all RDD function operations, analyze the operation logic between RDDs, and obtain the triple t consisting of all input RDDs, dependencies, and output RDDs:
[0040] t=R RhR ⊕r⊕R Rt
[0041] Where R RhR is the input RDD ID, r is the dependency, R RtR is the output RDD ID. At the same time, the RDD that cannot form a triple relationship is discarded.
[0042] In step 1.2, this embodiment: read all triples t, merge the same RDDs into the same node, and construct a directed acyclic graph DAG, wherein the RDD data that cannot form a triple is discarded.
[0043] Step 2: Filter the RDD data set, specifically:
[0044] In step 2.1, this embodiment sets a threshold for the storage capacity of the current memory, and once the storage capacity in the memory reaches the specified threshold, a filtering operation is performed.
[0045] In step 2.2, this embodiment: by assigning weights to each attribute of the RDD, the total weight value of the RDD partition is calculated. The formula for weight calculation is defined as follows:
[0046]
[0047] Among them, w represents the weight value of the RDD partition, Indicates the computational cost of the RDD. Indicates the number of times the RDD is used, S p Indicates the size of the partition. Indicates the life cycle of the RDD. Indicates the position of the RDD input RDD for calculation. The elements in A = {α0, α1, α2, α3, α4} are constants, respectively and The normalized weight of , the weight value is a hyperparameter, which is determined by the user's specific task requirements.
[0048] Step 4: perform iterative calculation, specifically:
[0049] In step 4.1, in this embodiment: a data set is allocated to each thread block in the global video memory and placed in the shared memory.
[0050] In step 4.2, this embodiment: let each thread in the thread block read the data to be calculated in the global video memory in turn, and perform iterative calculation;
[0051] In step 4.3, this embodiment: updates the data set stored in the shared memory with the data set iteratively calculated in step 4.2;
[0052] In step 4.4, in this embodiment: in order to improve the computing efficiency of the shared memory, a threshold is set for the storage capacity of the shared memory. When the storage capacity of the shared memory reaches a certain threshold, the data with low local access frequency in each iteration is transferred to the global video memory;
[0053] Step 5: perform incremental iterative calculation, specifically:
[0054] In step 5.1, in this embodiment: when the data set undergoes incremental changes, that is, when new iterative data is generated due to dynamic changes in the stream data, the shared memory reads the incremental data from the global video memory;
[0055] In step 5.2, this embodiment: by performing incremental iterative calculation, a data set after incremental iterative calculation is obtained. Incremental iteration is an iterative method for obtaining a new iterative result based on incremental data and original iterative results, and the incremental data is new iterative data generated due to business growth.
[0056] In step 5.3, this embodiment uses the data set calculated by the incremental iteration in step 5.2 to update the data set in the shared memory. The data whose loss values are all less than the specified threshold for three iterations are determined as converged data, and the converged data is stored in the global memory and used in subsequent iterations without being updated, so as to achieve the purpose of incremental control.
[0057] In step 5.4, in this embodiment: after the shared memory reaches a certain threshold, the data that is not frequently accessed is moved to the global video memory.
[0058] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An incremental iterative method for distributed deep learning, characterized in that: The following steps are involved: S1. Construct DAG; S1.
1. Input sample data, traverse all RDD function operations, and obtain the triple t consisting of all input RDDs, dependencies, and output RDDs: t = R h ⊕r⊕R t Where R h is the input RDD ID, r is the dependency, R t It is the output RDD ID. At the same time, the RDD that cannot form a triple relationship is discarded. S1.2, read all triples t and build DAG; S2, filter the RDD data set; S2.1, determine whether the current memory storage capacity reaches the threshold, if reached, continue with the next step; S2.2, calculate the RDD weight; the formula of this process is defined as follows: ; Among them, w represents the weight value of the RDD partition, Indicates the computational cost of the RDD. Indicates the number of times the RDD is used. Indicates the size of the partition. Indicates the life cycle of the RDD. Indicates the location of the input RDD for calculating this RDD; are constants, respectively The weight value is determined by the user's specific task requirements; S3, convert the filtered RDD data set into the data type that can be processed by the GPU and store it in the GPU global memory; S4, performing iterative calculation; S4.1, allocate a data set for each thread block in the global video memory and place it in the shared memory; S4.2, each thread in the thread block reads the data to be calculated in the global video memory in turn, and performs iterative calculation; S4.3, updating the data set stored in the shared memory with the data set iteratively calculated in step S4.2; S4.4, when the shared memory reaches a certain threshold, the data with low local access frequency in each iteration is transferred to the global video memory; S5, performing incremental iterative calculation; S5.1, when the data set undergoes incremental changes, that is, when new iterative data is generated due to dynamic changes in the stream data, the shared memory reads the incremental data from the global video memory; S5.
2. Perform incremental iterative calculation to obtain a data set after incremental iterative calculation; wherein incremental iteration is an iterative method for obtaining a new iterative result based on incremental data and an original iterative result, and the incremental data is new iterative data generated due to business growth; S5.3, updating the data set in the shared memory with the data set calculated by the incremental iteration in step S5.2 above; wherein the data judged to be converged is stored in the global memory and is continued to be used in subsequent iterations without being updated, so as to achieve the purpose of incremental control; S5.
4. When the shared memory reaches a certain threshold, the data that is not frequently accessed is moved to the global video memory.
Citation Information
Patent Citations
Caching optimizing method of internal storage calculation
CN103631730A
Classification method of Stage based on resilient distributed dataset (RDD) and terminal
CN106339458A