Optimization Method for Parallel Deep Convolutional Neural Network Based on Im2col

By adopting the parallel deep convolutional neural network optimization method based on Im2col in a big data environment, combined with MHO-PFES, IM-PMTS and IM-BGDS strategies, the problems of slow convolutional layer operation speed and poor convergence of loss function in DCNN model training are solved, and efficient model training and accurate results are achieved.

CN114819136BActive Publication Date: 2025-06-13SHAOGUAN COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210279825.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-21
Publication Date
2025-06-13
Estimated Expiration
2042-03-21

AI Technical Summary

Technical Problem

In the big data environment, DCNN model training faces the problems of slow convolutional layer operation speed and poor convergence of loss function.

Method used

A parallel deep convolutional neural network optimization method based on Im2col is proposed. Through the steps of parallel feature extraction, parallel model training and parallel parameter update, MHO-PFES, IM-PMTS and IM-BGDS strategies are adopted to solve the problems of data redundant features, convolutional layer operation speed and poor convergence of loss function, respectively.

Benefits of technology

It effectively avoids the problem of many redundant features of data, improves the operation speed of convolutional layer, solves the problem of poor convergence of loss functions, and significantly improves the operation efficiency of the algorithm and the accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114819136B_ABST
    Figure CN114819136B_ABST
Patent Text Reader

Abstract

The present invention proposes an optimization method for a parallel deep convolutional neural network based on Im2col, comprising the following steps: S1, feature parallel extraction: extracting target features in data as the input of the convolutional neural network; S2, model parallel training: in the convolutional process during the parallel DCNN model training stage, completing distributed convolutional kernel pruning and multi-node convolutional calculation through the IM-PMTS strategy; and combining the MapReduce and Im2col methods to parallelly train the model; S3, parameter parallel update: in the backpropagation stage, adopting the IM-BGDS strategy to update parameters for batch data; S4, inputting the data to be measured into the DCNN model after parameter parallel update and outputting the classification result. The MHO-PFES strategy proposed by the present invention can avoid the problem of redundant data features; the IM-PMTS strategy improves the operation speed of the convolutional layer; the IM-BGDS strategy eliminates the influence of abnormal data on the batch gradient and solves the problem of poor convergence of the loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data mining, and particularly to an optimization method for a parallel deep convolutional neural network based on Im2col. Background Art

[0002] As an important classification algorithm in the field of deep learning, DCNN has powerful representation ability, generalization ability and fitting ability, with stable effects and no need for additional feature engineering on data. It is often used in fields such as image classification, speech recognition, object detection, semantic segmentation, face recognition, and autonomous driving, and has received extensive attention and in-depth research.

[0003] With the rapid development of Internet technology and the advent of the big data era, big data has the "4V" characteristics of large volume, high velocity, variety, and high value compared with traditional data. The "4V" characteristics lead to difficulties such as a large amount of time consumption brought by training the DCNN model with massive data, and the need for repeated training of model parameters due to data and modality changes. Therefore, how to reduce the training cost of the DCNN model in a big data environment has become an urgent problem to be solved.

[0004] In recent years, the MapReduce parallel computing model developed by Google has been favored by many scholars and enterprises for its advantages such as easy programming, high fault tolerance, balanced load, and strong scalability. Many DCNN algorithms based on the MapReduce computing model have also been widely studied. Leung J et al. proposed a parallel DCNN algorithm based on MapReduce. This algorithm adopts the idea of divide and conquer, divides the data through the Split method of MapReduce, constructs multiple computing nodes to train the DCNN network model simultaneously, and selects the network model with the highest accuracy as the output of the algorithm, realizing the parallel training process of DCNN. Based on this, Huang X et al. proposed the parallel deep convolutional neural network algorithm FCNN (Fully CNN for processing CT scan image). The algorithm transforms the full view into a sparse view and smooths the feature edges through a Gaussian filter to enhance important texture feature information. Although the algorithm can speed up the reading speed during the process of transforming the full view into a sparse view, due to the change in the feature structure of the sparse view, it is difficult to screen features, resulting in a problem of many redundant feature data during the training of the model. Wang H et al. proposed the single-stride optimized CNN algorithm SSOCNN (An optimization of im2col, an important method of CNNs based on continuous address access) based on the Im2col method. This algorithm designs an acceleration method for the im2col algorithm in the case of a single stride based on continuous memory address reading, accelerates the process of mapping an image into a matrix by changing the data reading order, and uses general matrix multiplication to perform matrix multiplication operations on column vectors and convolutional kernels, realizing the acceleration of convolutional layer operations. However, in the process of constructing parallel convolutional operations, the algorithm is difficult to eliminate redundant convolutional kernels scattered in each node, resulting in the inability to solve the problem of slow convolutional layer operation speed in a large data environment. Mao et al. combined DCNN with the firefly algorithm to propose the MR-FPDCNN algorithm (Deep convolutional neural network algorithm based on feature graph and parallel computing entropy using MapReduce). This algorithm combines the information sharing search strategy with the firefly algorithm to find the optimal parameters of the network model and shares the DCNN network parameters through the MapReduce communication mechanism, accelerating the convergence speed of the loss function. However, the firefly algorithm has poor robustness. When dealing with abnormal data (such as mislabeled, noisy data, etc.), it will cause the convergence of the loss function to oscillate, resulting in poor convergence of the loss function. Summary of the Invention

[0005] The present invention aims to solve at least the technical problems existing in the prior art, and particularly innovatively proposes an optimization method for a parallel deep convolutional neural network based on Im2col.

[0006] In order to achieve the above object of the present invention, the present invention provides an optimization method for a parallel deep convolutional neural network based on Im2col, including the following steps:

[0007] S1, Feature parallel extraction: Extract the target features in the data as the input of the convolutional neural network, solving the problem of a large number of redundant features in the data;

[0008] S2, Model parallel training: In the convolutional process of the parallel DCNN model training stage, complete distributed convolutional kernel pruning and multi-node convolutional calculation through the IM-PMTS strategy; and combine the MapReduce and Im2col methods to parallel train the model, improving the operation speed of the convolutional layer;

[0009] S3, Parameter parallel update: In the backpropagation stage, adopt the IM-BGDS strategy to update the parameters for batch data. This strategy for batch data can exclude the gradient descent method of abnormal data points and avoid the influence of abnormal data points on the gradient of the batch data.

[0010] S4, Input the data to be tested into the DCNN model after parameter parallel update, and output the classification result.

[0011] Further, the S1 adopts the MHO-PFES strategy for feature parallel extraction, and the MHO-PFES strategy includes the following steps:

[0012] S1-1, Feature extraction: Use an improved non-mean filter to filter the input data, calculate the Laplace equation h(x, y) of the filtered data, and find the zero crossings of the Laplace equation to extract data features;

[0013] S1-2, Feature screening: To further screen the target features, propose a feature correlation index FCI(u, v) to compare the similarity between any two data blocks, and set a correlation coefficient ε. Reduce the redundant features in the data by removing the data blocks where FCI(u, v) < ε.

[0014] Further, the improved non-mean filter FT(a, b) includes:

[0015]

[0016] where a represents the target window matrix;

[0017] b represents the neighborhood window matrix;

[0018] θ(·) is the feature transformation function;

[0019] G i is the current data;

[0020] are the vectorized representations of matrices a and b respectively;

[0021] |·| represents the modulus of a vector.

[0022] Furthermore, the feature correlation index FCI(u, v) includes:

[0023]

[0024] where μ u , μ v represent the expectations of u and v respectively;

[0025] σ u , σ v represent the variances of u and v respectively;

[0026] u and v represent two feature vectors respectively.

[0027] Furthermore, the IM-PMTS strategy in S2 includes the following steps:

[0028] S2-1, Convolution kernel pruning: Design the Mahalanobis distance center value MDCV, find the vector linearly related to the convolution kernel in the network model by solving the MDCV value, calculate the distance dist between this vector and each convolution kernel, and prune the convolution kernel with dist < α by setting the threshold α to reduce the redundant parameters in the network model;

[0029] S2-2, Parallel Im2col convolution: Use the Im2col algorithm to map the feature map into a matrix, store the key-value pairs of the matrix and the corresponding convolution kernel, distribute them to each computing node for matrix operations to accelerate the operation of the convolution layer, obtain the operation result of the convolution layer operation, and store the result in HDFS.

[0030] Furthermore, the Mahalanobis distance center value MDCV includes:

[0031]

[0032] where μ represents the mean of all convolution kernels;

[0033] S represents the covariance matrix of all convolution kernels;

[0034] R n is the set of convolution kernels in the same hierarchical model, R n ={X1 , X 2 ,..., X n}, x ∈ R n , x takes any value from {X 1 , X 2 ,..., X n}, where X 1 , X 2 ,..., X n represents the convolutional kernel in the network model;

[0035] T represents the transpose.

[0036] Further, the IM - BGDS strategy includes the following steps:

[0037] S3 - 1, Gradient construction: Propose the loss mean weight LAW(g i ) to exclude the influence of abnormal data on the batch gradient, and design the loss summation gradient LSG(T) to construct the average gradient of batch data, solving the problem of poor convergence of the loss function;

[0038] S3 - 2, Parameter parallel update: After obtaining the average gradient of the batch data, combine the MapReduce computing framework and the error conduction formula of backpropagation to calculate the error in parallel, realizing the parallel update of parameters.

[0039] Further, the loss mean weight LAW(g i ) includes:

[0040]

[0041] Where:

[0042]

[0043] Where LAD(g i ) is the absolute value of the difference between the loss function value of data g i and the mean of the loss function values;

[0044] g i represents a piece of data in the batch data;

[0045] τ is the threshold for measuring LAD(g i );

[0046] batch_size represents the batch data size;

[0047] J(ω, b) i represents the loss function value of data g i ;

[0048] ω, b are the convolutional kernel parameters and the bias of the convolutional layer respectively.

[0049] Furthermore, the loss summation gradient LSG(T) includes:

[0050]

[0051] where batch_size represents the batch data size;

[0052] denotes the gradient of the loss function of data g i with respect to parameter x;

[0053] T represents all data in the batch;

[0054] LAW(g i ) is the weight index of the loss function value of data g i .

[0055] In summary, due to the adoption of the above technical solutions, the MHO-PFES strategy proposed by the present invention can avoid the problem of excessive redundant features in data; the IM-PMTS strategy improves the operation speed of the convolutional layer; the IM-BGDS strategy eliminates the influence of abnormal data on the batch gradient and solves the problem of poor convergence of the loss function.

[0056] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0058] Figure 1 is the speedup ratio of each algorithm on the CIFAR10 and ImageNet1K datasets, where Figure 1 (a) is the speedup ratio of each algorithm on the CIFAR10 dataset, Figure 1 (b) is the speedup ratio of each algorithm on the ImageNet1K dataset.

[0059] Figure 2 is the Top-1 accuracy of each algorithm on the CIFAR10 and ImageNet1K datasets, where Figure 2 (a) is the Top-1 accuracy of each algorithm on the CIFAR10 dataset, Figure 2 (b) is the Top-1 accuracy of each algorithm on the ImageNet1K dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0061] The present invention proposes an optimization method for a parallel deep convolutional neural network based on Im2col. The specific embodiments are as follows and include the following steps:

[0062] S1, Feature parallel extraction: Extract the target features in the medical image data as the input of the convolutional neural network;

[0063] S2, Model parallel training: During the convolution process in the parallel DCNN model training stage, complete distributed convolutional kernel pruning and multi-node convolution calculation through the IM-PMTS strategy; and combine the MapReduce and Im2col methods to train the model in parallel;

[0064] S3, Parameter parallel update: In the backpropagation stage, use the IM-BGDS strategy to update the parameters for the batch of medical image data;

[0065] S4, Input the medical image test data into the DCNN model after parameter parallel update, and output the classification result of the medical image.

[0066] Based on the advantages of the MapReduce programming model, this invention proposes an optimized algorithm for parallel deep convolutional neural networks (IA-PDCNNOA) based on the Im2col algorithm. First, a parallel feature extraction strategy MHO-PFES (Parallel feature extraction strategy based on Marr Hildreth operator) based on the Marr-Hildreth operator is proposed to extract target features in the data as the input of the convolutional neural network, effectively avoiding the problem of excessive redundant features in the data. Second, a parallel model training strategy IM-PMTS (Parallel model training strategy based on Im2col method) based on the Im2col method is designed. By designing the Mahalanobis distance center value to remove redundant convolutional kernels and combining the MapReduce and Im2col methods to train the model in parallel, the operation speed of the convolutional layer is improved. Finally, an improved mini-batch gradient descent strategy IM-BGDS (Improved Mini Batch gradient descent strategy) is proposed to eliminate the influence of abnormal data on the batch gradient and solve the problem of poor convergence of the loss function. The algorithm proposed in this invention has significantly improved both in terms of running efficiency and model accuracy. In addition, the knowledge mined through this method can provide great help in biology, medicine, astronomy, and geography.

[0067] 1. Parallel Feature Extraction

[0068] Currently, in the parallel DCNN algorithm in the big data environment, there is a problem of excessive redundant features during the model training process. To solve this problem, the MHO-PFES strategy based on the Marr-Hildreth operator is proposed. This strategy mainly includes two steps: (1) Feature extraction: An improved non-mean filter FT(a,b) (Filter transformation) is proposed to filter the input data, and the Laplace equation h(x,y) of the filtered data is calculated. The zero crossings of the Laplace equation are found to extract data features; (2) Feature screening: To further screen the target features, a feature correlation index FCI(u,v) (Feature correlation indices) is proposed to compare the similarity between any two data blocks, and a correlation coefficient ε is set. By removing the data blocks with FCI(u,v) < ε, the redundant features in the data are reduced.

[0069] 1.1 Feature Extraction

[0070] In order to obtain high-precision data features, it is necessary to first remove noise from the initial dataset. Therefore, a non-mean filter FT(a,b) based on cosine similarity is proposed to remove data noise through the self-similarity of data in different regions. Then, the Laplacian operation of the convolution kernel f(x,y) and the data g(x,y) is used to construct and find the zero-crossing of the Laplace equation to extract data features. The specific process is as follows: First, set the target window matrix a and the neighborhood window matrix b, and slide the neighborhood window in the current data. By comparing the cosine similarity of matrices a and b, the weighted value of the neighborhood window is obtained, and the data is denoised according to the weight value and the gray value of each point itself to obtain the denoised image g(x,y). Then, set the convolution kernel f(x,y) with a size of 3*3 and perform the Laplacian operation on g(x,y) to obtain the Laplace equation where x and y respectively represent the pixel values of the image at (x,y), is the Laplacian operator, a represents the target window matrix, and b represents the neighborhood window matrix. Finally, judge whether the second derivative of the current node is a cross zero, and the first derivative of this node is at a large peak. If the condition is satisfied, this node is retained; otherwise, this pixel point is set to zero, and then the current data nodes are merged to obtain the data after feature extraction. Generally speaking, for the non-mean denoising algorithm, the data refers to image data.

[0071] Theorem 1 (Non-mean filter FT(a,b) based on cosine similarity): Given that a represents the target window matrix, b represents the neighborhood window matrix, a, b ∈ (x,y), and (x,y) represents the current data. The calculation formula of the transformation function FT(a,b) is as follows:

[0072]

[0073] where θ(·) is the feature transformation function, which can be, for example, a linear kernel function, a Gaussian kernel function, etc.; G i is the current data, are the vectorized representations of matrices a and b respectively, and |·| represents the modulus of the vector.

[0074] Proof: The non-local mean filtering principle utilizes the non-correlation feature of noise. Let the value of the noise-free pixel block be ω(p,q) and the noise value be ψ(p,q). Then the value of the pixel block after fusion with noise is ρ(p,q) = ω(p,q) + ψ(p,q). The mean value is obtained by superimposing similar pixel blocks and taking the mean where ρi(p,q) represents the pixel value of the i-th pixel block after fusion with noise, and k is the total number of pixel blocks. Then The expectation of is Due to the similarity of pixel blocks, E[ω i(p,q) can be simplified to ω(p,q). When the noise is 0, E[ψ(p,q)] = 0, so In addition, due to the non - correlation of the noise, the variance of ω(p,q) is Since ω(p,q) has no noise and its variance is 0, so It indicates that the noise ψ(p,q) is related to the variance, and FT(p,q) reduces the data noise by reducing ψ(p,q). Q.E.D.

[0075] 1.2 Feature Screening

[0076] After feature extraction is completed, the strategy divides the data in the batch into chunks, and proposes a feature - correlation index FCI(u,v) to calculate the feature similarity between any two data chunks. Then, it removes the data chunks with FCI(u,v) < ε to achieve the removal of redundant features in the data. The specific process is as follows: First, divide the data of the same category into the batch, divide the data in the batch into equal - sized data chunks, and number each data chunk in order. Calculate the feature - correlation index FCI(u,v) between any two data chunks, and store the key - value pair <(u,v),FCI(u,v)> in HDFS; then, set the correlation coefficient ε, and traverse the key - value pair <(u,v),FCI(u,v)> in order to remove the items with FCI(u,v) < ε; finally, traverse the key - value pair <(u,v),FCI(u,v)> again, read the key values of all key - value pairs to obtain the subscripts of the target feature data chunks, and splice the screened data chunks to obtain the input data of the convolutional neural network, completing the feature screening of the data.

[0077] Theorem 2 (Feature - correlation index FCI(u,v)): Given that u and v represent two feature vectors respectively, μ u , μ v represent the expectations of u and v, and σ u , σ v represent the variances of u and v. The calculation formula of the feature - correlation index FCI(u,v) is as follows:

[0078]

[0079] Proof: FCI(u,v) is an index to measure the feature similarity between u and v. Let μ u , μ v represent the expectations of u and v, and σ u , σ v represent the variances of u and v. When the feature vector u has σ u = 0, the operation of the convolution process on u belongs to linear superposition and cannot extract features. At this time, FCI(u,v) = 0; when σ u ≠0, σ vWhen it is not equal to 0 and the eigenvectors x and y are eigen-similar, FCI(u, v) approaches 1, where → means approaching. Q.E.D.

[0080] 2. Model parallel training

[0081] In the current DCNN algorithm in the big data environment, parallel training of the model requires distributing feature maps and convolutional kernels to different computing nodes for operations. However, during the construction of parallel convolutional operations, it is difficult for the algorithm to filter out redundant convolutional kernels scattered across various nodes, resulting in the inability to solve the problem of slow operation speed in the convolutional layer in the big data environment. To solve this problem, this paper proposes the IM-PMTS strategy, which mainly consists of two steps: (1) Convolutional kernel pruning: Design the Mahalanobis distance center value (MDCV), find vectors linearly related to the convolutional kernels in the network model by solving the MDCV value, calculate the distance dist between this vector and each convolutional kernel, and reduce redundant parameters in the network model by pruning convolutional kernels with dist < α by setting a threshold α; (2) Parallel Im2col convolution: Use the Im2col algorithm to map the feature map into a matrix, store the key-value pairs of the matrix and the corresponding convolutional kernels, distribute them to each computing node for matrix operations to accelerate the operations in the convolutional layer, obtain the operation results of the convolutional layer, and store the results in HDFS, i.e., the Hadoop Distributed File System.

[0082] 2.1 Convolutional kernel pruning

[0083] To reduce the invalid calculations caused by redundant convolutional kernels in the convolutional neural network, design the Mahalanobis distance center value (MDCV) to filter out redundant convolutional kernels in the current convolutional layer, thereby accelerating the operations in the convolutional layer. The specific process is as follows: First, calculate the covariance matrix S and mean μ of all convolutional kernels X 1 , X 2 ,..., X n in the convolutional layer, and construct the objective function f(x) of MDCV; then, calculate the second-order Taylor expansion of f(x) at its stationary point x k . Here, represents the Laplace operator, and (·) T represents the transpose; if the current second derivative is non-singular, the next iteration point is If the current second derivative is singular, first solve to determine the search direction d k , and then determine the next iteration point x k+1 = x k + d k, until the optimal MDCV value is found; finally, calculate the distance dist from all convolution kernels in the convolutional layer to the MDCV value, and set a threshold α to prune the convolution kernels with dist < α to complete the convolution kernel pruning process. Here, k is the number of search times.

[0084] Theorem 3 (Mahalanobis Distance Center Value MDCV): Given X 1 , X 2 ,..., X n represent the convolution kernels in the network model, S represents the covariance matrix of all convolution kernels, and μ represents the mean of all convolution kernels. The calculation formula for the Mahalanobis Distance Center Value MDCV is as follows:

[0085]

[0086] where Rn is the set of convolution kernels in the same hierarchical model, and T represents the transpose.

[0087] Proof: MDCV is the minimum distance from the eigenvector x to the eigenvector group X 1 , X 2 ,..., X n Let S be the covariance matrix of the vector group X 1 , X 2 ,..., X n , and μ be the mean of the vector group. The covariance matrix S is introduced to exclude the interference of the correlation between variables. When the eigenvector x → MDCV value, the eigenvector x is more easily replaced by the eigenvector group. When x = MDCV, x is linearly correlated with X 1 , X 2 ,..., X n . Therefore, the MDCV value represents the minimum distance from the eigenvector x * to the eigenvector group X 1 , X 2 ,..., X n . Q.E.D.

[0088] 2.2 Parallel Im2col Convolution

[0089] After completing the convolution kernel pruning, the parallel operation of Im2col convolution can be realized by combining the MapReduce computing framework. The specific process is as follows: First, map the input feature map M i to the convolution calculation matrix I i , and store the key-value pair <I i , K i , K z > for each mapped matrix I z , where K iThe corresponding convolutional kernels have a many-to-many relationship; then, the Map() function is called to matrix multiply the matrix I in the key-value pair i with the one-dimensional vector of the corresponding convolutional kernel to obtain the intermediate convolution result; finally, the Reduce() function is called to merge the feature maps of the same data to obtain the final output feature map NM i .

[0090] 3. Parallel Parameter Update

[0091] In the current parallel DCNN algorithm under big data, the random gradient descent method or the batch gradient descent method is used to update the parameters during the backpropagation process. However, in the process of implementing gradient descent, the training of the DCNN model on abnormal data (mislabeled, noisy data, etc.) will cause the loss function to converge and oscillate, resulting in poor convergence of the loss function. To solve this problem, the IM-BGDS strategy is proposed, which mainly includes two steps: (1) Gradient construction: The loss average weight LAW(g i )(LossAverage Weight) is proposed to exclude the influence of abnormal data on the batch gradient, and the loss sum gradient LSG(T)(Loss SumGradient) is designed to construct the average gradient of the batch data, solving the problem of poor convergence of the loss function; (2) Parallel parameter update: After obtaining the average gradient of the batch data, the MapReduce computing framework and the error conduction formula of backpropagation are combined to calculate the error in parallel, realizing the parallel update of the parameters.

[0092] (1) Gradient construction

[0093] To exclude the influence of abnormal data on the batch gradient, the loss average weight LAW(g i ) and the loss sum gradient LSG(T) are designed to solve the problem of poor convergence of the loss function. The specific process is as follows: First, when updating the parameters, calculate the mean value of the loss function of the entire batch of data, and subtract the mean value from the loss function value of each data g i to construct the loss average weight LAW(g i ), and store the key-value pair <g i ,LAW(g i )> in HDFS; then, calculate the partial derivative of the loss function of each data g i with respect to the current parameter δ z Store the key-value pair in HDFS, and set batch_size to the number of 1s in LAW(g i ); finally, traverse the key-value pair <g i with g i as the index,LAW(g i )> and​ Construct the average gradient LSG(T) of the batch data to obtain the batch gradient of the current parameters.

[0094] Theorem 4 (Loss Mean Weight LAW(g i )): Given g i represents a piece of data in the batch data, J(ω, b) i represents the loss function value of the data g i ω and b are the convolution kernel parameters and the bias of the convolutional layer respectively, batch_size represents the batch data size, and LAD(g i ) is the absolute value of the difference between the loss function value of the data g i and the mean of the loss function values. The calculation formula of the loss mean weight LAW(g i ) is as follows:

[0095]

[0096] Where:

[0097]

[0098] Proof: LAW(g i ) is the weight index of the loss function value of the data g i . Let batch_size be the batch data size, and τ be the threshold for measuring LAD(g i ). When LAD(g i ) < τ, the loss function value of the current data g i belongs to the normal value, so let LAW(g i ) = 1 to keep it; when LAD(g i ) ≥ τ, the loss function value of the current data g i belongs to the outlier value, so let LAW(g i ) = 0. Q.E.D.

[0099] Theorem 5 (Loss Summation Gradient LSG(T)): Given that T represents all the data in the batch, represents the gradient of the loss function of the data g i with respect to the parameter x, and batch_size represents the batch data size. The calculation formula of the loss summation gradient LSG(T) is as follows:

[0100]

[0101] Proof: LSG(T) is the average gradient of the batch data batch. Let be the gradient of the loss function of the data g i with respect to the parameter x, and batch_size be the batch data size. When LIW(g i) = 1, the gradient of data g i descends towards the optimal direction; when LIW(g ) = 0, the gradient of data g i has a large deviation from the optimal direction and is not included in the LSG(T) gradient. Q.E.D. i The gradient of

[0102] (2) Parameter parallel update

[0103] After obtaining the average gradient of the batch data, use the error backpropagation algorithm to parallelize the update of the error term parameters, and combine the MapReduce computing framework to obtain the network model after parameter parallel update. The parameter parallel update process is as follows: First, calculate the gradients of all parameters of the convolutional kernel in the l-1 layer and map the results to key-value pairs and store them in HDFS; then, calculate the change amount of the parameters of the convolutional kernel in the network model to update the network parameters of the convolutional kernel in the l-1 layer, where r is the convolutional kernel number, and its function is to correspond to the corresponding gradient. Finally, synchronize the updated parameters to all computing nodes through HDFS and perform the next update until all parameters in the network model are updated. The value range of l depends on the number of convolutional layers of the network model adopted.

[0104] 4. Effectiveness of the Parallel Deep Convolutional Neural Network Optimization Algorithm Based on Im2col (IA-PDCNNOA)

[0105] To verify the performance effect of the algorithm IA-PDCNNOA, we applied the IA-PDCNNOA method to two datasets, ImageNet1K dataset and CIFAR10, and their specific information is shown in Table 1. The MR-FPDCNN, SSOCNN, and FCNN algorithms were compared in terms of algorithm parallel performance, classification accuracy, etc.

[0106] Table 1 Dataset detailed information

[0107] Items CIFAR10 ImageNet 1K Number of pictures / sheets 60 000 1281 167 Picture size / pixel 32*32 224*224 Number of categories / categories 10 1000

[0108] 4.1 Experimental analysis of the speedup ratio of the IA-PDCNNOA algorithm

[0109] ​​​To verify the parallel performance of the IA-PDCNNOA algorithm in a big data environment, based on the CIFAR10 and ImageNet 1K datasets, this paper uses the speedup as a measurement metric and compares it with the MR-FPDCNN, SSOCNN, and FCNN algorithms. Meanwhile, to ensure the accuracy of the experimental results, the average running time of each algorithm for 10 runs is taken to calculate the speedup, which is used as the final experimental result. The experimental results are as Figure 1 shown:

[0110] As can be seen from Figure 1 (a), when dealing with a relatively small-scale dataset like CIFAR10, the speedup of each algorithm increases slowly with the increase in the number of nodes. Among them, when the number of cluster nodes is 4, the speedup of IA-PDCNNOA is 0.3 and 0.5 lower than those of the FCNN and SSOCNN algorithms with low parallelization degrees, respectively. However, in Figure 1 (b), when the algorithm processes the relatively large ImageNet 1K dataset, the speedup of the IA-PDCNNOA algorithm increases rapidly, reaching 9.8 when the number of cluster nodes is 8, which is 1.1, 4.1, and 4.6 higher than those of the ME-FPDCNN, FCNN, and SSOCNN algorithms, respectively. The reasons for these results are as follows: when the IA-PDCNNOA algorithm processes a relatively small-scale dataset, the distribution of data to each computing node will cause the communication time overhead between nodes to increase rapidly, and the improvement in the running speed obtained through parallel computing is extremely limited; when the IA-PDCNNOA algorithm processes a relatively large-scale dataset, due to its designed IM-PMTS strategy, by proposing the Mahalanobis distance center value MDCV to prune the convolutional kernels of the same layer, the overhead of convolutional layer parameters in network communication is reduced, and then by combining the MapReduce and Im2col methods for parallel training to accelerate the convolutional operation process, the operation speed of the convolutional layer is improved, and the speedup of the algorithm is increased. The experiment shows that the parallelization ability of the IA-PDCNNOA algorithm is significantly enhanced with the increase in the number of cluster nodes. It is suitable for parallel processing of large datasets and has good performance.

[0111] 4.2 Experimental Analysis of the Accuracy of the IA-PDCNNOA Algorithm

[0112] To further verify the training effect of the IA-PDCNNOA algorithm, the Top-1 accuracy is used as a measurement metric to evaluate the training effect of the algorithm. IA-PDCNNOA, MR-FPDCNN, SSOCNN, and FCNN are respectively processed on the CIFAR10 and ImageNet1K datasets, and their Top-1 accuracy is calculated as the experimental result. The experimental results are as Figure 2 shown:

[0113] From Figure 2 (a), it can be seen that when dealing with a relatively small-scale dataset such as CIFAR10, the Top-1 accuracy of each algorithm can be stably maintained at a relatively high value. Among them, the IA-PDCNNOA algorithm has the highest Top-1 accuracy and converges earlier, reaching 89.72%. Compared with the MR-FPDCNN, SSOCNN, and FCNN algorithms, it is 2.87%, 4.62%, and 6.48% higher. However, in Figure 2 (b), when the algorithm processes the relatively large ImageNet 1K dataset, there are significant differences in the Top-1 accuracy and algorithm convergence of each algorithm. Among them, the IA-PDCNNOA algorithm has the highest Top-1 accuracy among the four parallelization algorithms, reaching 72.41%. Compared with the MR-FPDCNN, SSOCNN, and FCNN algorithms, it is 2.31%, 7.98%, and 2.85% higher. However, the other three algorithms all have varying degrees of difficulty in converging. These results are because the IA-PDCNNOA algorithm proposes the IM-BGDS strategy, which designs the loss summation gradient LSG(T) to construct the mini-batch data gradient and updates the parameters in parallel through the error backpropagation algorithm, excluding the influence of abnormal data on the batch gradient and enhancing the convergence of the IA-PDCNNOA algorithm. Experimental data shows that IA-PDCNNOA has a higher convergence speed and accuracy compared to the other three parallelization algorithms, and it is suitable for the model parallelization training of deep convolutional neural networks under large datasets.

[0114] 4.3 Experimental Analysis of the Running Time and FLOPs of the IA-PDCNNOA Algorithm

[0115] To verify the algorithm execution speed and model optimization effect of the IA-PDCNNOA algorithm in a large data environment, this paper calculates the running time and FLOPs of Baseline, IA-PDCNNOA, MR-FPDCNN, SSOCNN, and FCNN based on the CIFAR10 and ImageNet 1K datasets. Among them, Baseline is the benchmark data of the ResNet50 model under a 1 / 8 data load. The experimental results are shown in Table 2:

[0116] Table 2 Running Time and FLOPs of Each Algorithm on Two Datasets

[0117]

[0118]

[0119] As can be seen from Table 2, when dealing with a relatively small-scale dataset such as CIFAR10, there is not much difference in the running time of each algorithm, but their floating-point operation amounts are all reduced to varying degrees. Among them, the floating-point operation amount of IA-PDCNNOA is reduced by 5%, 21%, and 16% compared with the MR-FPDCNN, SSOCNN, and FCNN algorithms respectively. However, when dealing with a larger dataset such as ImageNet1K, both the running time and the floating-point operation amount of the IA-PDCNNOA algorithm are better than the other three algorithms. Among them, the running time of the IA-PDCNNOA algorithm is 1.32×10 4 s, 3.85×10 4 s, and 5.29×10 4 s faster than the MR-FPDCNN, SSOCNN, and FCNN algorithms respectively, and the floating-point operation amounts are reduced by 3%, 13%, and 8% respectively. These results are due to the MHO-PFES strategy proposed by the IA-PDCNNOA algorithm. By proposing the feature correlation index FCI(u,v), it removes the redundant features in the data and screens the target features of the data as the input of the convolutional neural network, reducing the floating-point operation amount of the model and accelerating the running speed of the algorithm. Generally speaking, by comparing the changing trends of the running time and floating-point operation amounts of the four algorithms on CIFAR10 and ImageNet 1K, it can be seen that as the training dataset increases, the reduction in the running time and floating-point operation amount of the IA-PDCNNOA algorithm has a large gap with other algorithms. Therefore, it can be concluded that IA-PDCNNOA is superior to MR-FPDCNN, SSOCNN, and FCNN and is suitable for the parallel training of DCNN models under large datasets.

[0120] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. An optimization method for a parallel deep convolutional neural network based on Im2col, characterized in that, it includes the following steps: S1, Feature parallel extraction: Extract the target features in the medical image data as the input of the convolutional neural network; S2, Model parallel training: During the convolution process in the parallel DCNN model training stage, complete distributed convolutional kernel pruning and multi-node convolution calculation through the IM-PMTS strategy; and combine the MapReduce and Im2col methods to parallel train the model; S3, Parameter parallel update: In the backpropagation stage, use the IM-BGDS strategy to update the parameters for batch medical image data; S4, Input the test medical image data into the DCNN model after parameter parallel update, and output the classification result of the medical image; The S1 uses the MHO-PFES strategy for feature parallel extraction, and the MHO-PFES strategy includes the following steps: S1-1, Feature extraction: Use an improved non-mean filter to filter the input data, calculate the Laplace equation h(x,y) of the filtered data, and find the zero crossings of the Laplace equation to extract data features; S1-2, Feature screening: To further screen the target features, propose a feature correlation index FCI(u,v) to compare the similarity between any two data blocks, and set a correlation coefficient ε, and reduce the redundant features in the data by removing the data blocks where FCI(u,v) < ε; The feature correlation index FCI(u,v) includes: where μ u , μ v represent the expectations of u and v respectively; σ u , σ v represent the variances of u and v respectively; u and v respectively represent two feature vectors; The IM-PMTS strategy in the S2 includes the following steps: S2-1, Convolutional kernel pruning: Design the Mahalanobis distance center value MDCV, find the vector linearly related to the convolutional kernel in the network model by solving the MDCV value, and calculate the distance dist between this vector and each convolutional kernel. By setting a threshold α, crop the convolutional kernels where dist < α to reduce the redundant parameters in the network model; S2-2, Parallel Im2col convolution: Use the Im2col algorithm to map the feature map into a matrix, store the matrix and the corresponding convolutional kernel key-value pairs, distribute them to each computing node for matrix operations to accelerate the operation of the convolutional layer, obtain the operation result of the convolutional layer operation, and store the result in HDFS; The Mahalanobis distance center value MDCV includes: where μ represents the mean of all convolutional kernels; S represents the covariance matrix of all convolutional kernels; R n is a set of convolution kernels in the same hierarchical model; T represents the transpose; The IM-BGDS strategy includes the following steps: S3-1, Gradient construction: Propose the loss mean weight LAW(g i ) to exclude the influence of abnormal data on the batch gradient, and design the loss summation gradient LSG(T) to construct the average gradient of batch data, solving the problem of poor convergence of the loss function; S3-2, Parameter parallel update: After obtaining the average gradient of the batch data, combine the MapReduce computing framework and the error conduction formula of backpropagation to calculate the error in parallel, and realize the parallel update of the parameters; The loss mean weight LAW(g i ) includes: where: where LAD(g i ) is the absolute value of the difference between the loss function value of data g i and the mean value of the loss function values; g i represents one piece of data in the batch data; τ is the threshold for measuring LAD(g i ); batch_size represents the batch data size; J(ω,b) i represents data g i the loss function value; ω and b are the convolutional kernel parameters and the bias of the convolutional layer respectively.

2. The optimization method for a parallel deep convolutional neural network based on Im2col according to claim 1, characterized in that, the improved non-mean filter FT(a,b) includes: where a represents the target window matrix; b represents the neighborhood window matrix; θ(·) is a feature transformation function; G i is the current data; They are respectively the vectorized representations of matrices a and b; |·| represents the norm of a vector.

3. According to the method for optimizing a parallel deep convolutional neural network based on Im2col as claimed in claim 1, characterized in that the loss summation gradient LSG(T) includes: where batch_size represents the batch data size; ▽J xi represents the gradient of the loss function for data g i with respect to parameter x; T represents all the data in the batch; LAW(g i ) is the weight index of the loss function value of the data g i .