Image classification method and device based on convolutional neural network and convolutional neural network

By adopting a convolutional neural network model with parallel structure in convolutional neural networks, using multiple parallel subconvolutional neural network layers for feature extraction and superimposing results, the problems of huge parameters and large calculations caused by increasing depth are solved, and the efficiency and accuracy of image classification are improved.

CN114091648BActive Publication Date: 2025-05-13CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010746150.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-29
Publication Date
2025-05-13
Estimated Expiration
2040-07-29

AI Technical Summary

Technical Problem

In the image classification task, existing convolutional neural networks have problems such as increasing depth, large calculation amount, saturation or degradation of recognition accuracy, and single structure type and few feature extraction.

Method used

A convolutional neural network-based image classification method is designed, and a convolutional neural network model with parallel structure is used to extract features through more than two parallel subconvolutional neural network layers, and the processing results are superimposed on the third dimension to reduce the number of parameters and calculation amount.

Benefits of technology

The number of eigenvalues ​​extracted is increased, the amount of calculations in the data processing process is reduced, the efficiency of data processing is improved, and the accuracy saturation or reduction caused by increasing depth is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114091648B_ABST
    Figure CN114091648B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses an image classification method, device, convolutional neural network, electronic device and computer-readable storage medium based on convolutional neural network. The method includes: obtaining first data, wherein the first data is represented by multiple images and category information corresponding to each image; training the convolutional neural network based on the first data; using the convolutional neural network trained based on the first data to identify the image to be identified, and obtaining category information corresponding to the image to be identified; the convolutional neural network includes a first processing layer, two or more parallel sub-convolutional neural network layers and a second processing layer connected in sequence; the present application proposes a brand-new convolutional neural network, which can increase the number of feature value extraction through two or more parallel sub-convolutional neural network layers, reduce the amount of calculation in the data processing process, and improve the efficiency of data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to image classification technology, and in particular to an image classification method, device, convolutional neural network, electronic device and computer storage medium based on convolutional neural network. Background Art

[0002] The current structure of convolutional neural networks is changing in two main directions. One is that the depth of the network is getting deeper and deeper. The classic convolutional neural network LeNet5 has only two convolutional layers and two sampling layers to extract features, and then connect the fully connected layer to achieve image classification. Although it has achieved a high recognition rate, there are still many problems. First of all, its success is based on a small data set. The image size is very small. The same training model cannot be applied to other larger data sets and other more complex recognition tasks. The computing performance of computers at that time was backward and could not meet the needs of large-scale computing. After a long time, convolutional neural networks did not achieve great success. In response to this problem, Krizhevsky et al. proposed a new convolutional neural network structure AlexNet. The structure of the entire network is wider and deeper. It successfully uses activation functions, regularization and data enhancement methods to apply convolutional neural networks to ImageNet, a large data set including 1.2 million images. It uses powerful GPU parallel computing and achieves the highest recognition rate compared to other algorithms. The sequential structure model VGGNet expands the depth of the convolutional neural network and uses smaller convolution kernels to stack convolutional layers, bringing the recognition rate of ImageNet to a better level.

[0003] The other is the change of structure. Increasing the depth of convolutional neural networks blindly will increase the number of parameters and the amount of calculation, making the model difficult to optimize. The number of parameters of AlexNet reached 60M, and increasing the depth requires higher computing performance of the computer. In 2013, the NIN structure was proposed. The parameters of the convolutional layer were reduced by a small 1x1 convolution kernel, and the global average pooling was proposed to address the problem that the parameters of the convolutional neural network were concentrated in the fully connected layer, which greatly reduced the parameters of the network. In the Inception module, different convolutional layers act on the same feature map, extracting diverse features and having strong generalization ability. Increasing the depth of the convolutional layer will lead to an increase in accuracy, but continuous increase will lead to unchanged or even decreased accuracy. To address this problem, the residual network ResNet uses the same strategy of continuously stacking small convolutional layers as VGGNet to increase the depth of the convolutional neural network, but uses residual modules. Although the depth of the network reaches 152 layers, the number of parameters is less than that of VGGNet, and the accuracy is higher.

[0004] As the requirements of recognition tasks become higher and higher, shallow convolutional neural networks can no longer meet the needs. At the same time, more and more researchers are optimizing convolutional neural networks, but there are still the following shortcomings:

[0005] The complexity and computational complexity of deep convolutional neural networks are relatively high;

[0006] As the depth of the convolutional neural network increases, the accuracy achieved is also getting higher and higher, but the parameters of the network are getting larger and larger. The recognition accuracy will first increase with the depth of the network, and then gradually reach saturation. After that, further increasing the depth of the network will cause the accuracy to remain unchanged or even decrease.

[0007] The convolutional neural network structure has a single structure type and extracts fewer features. As the scope of artificial intelligence applications continues to increase, the means of human-computer interaction are also becoming increasingly complex. Summary of the invention

[0008] The present application provides an image classification method, device, convolutional neural network, electronic device and computer storage medium based on convolutional neural network.

[0009] The embodiment of the present application provides an image classification method based on a convolutional neural network, wherein the convolutional neural network includes: a first processing layer, two or more parallel sub-convolutional neural network layers and a second processing layer; the output end of the first processing layer is respectively connected to the input end of the two or more parallel sub-convolutional neural network layers, and the output ends of the two or more parallel sub-convolutional neural network layers are respectively connected to the input end of the second processing layer; after the data input into the first processing layer is processed by the first processing layer, it can be processed by the two or more parallel sub-convolutional neural network layers respectively, and the two or more sub-processing results obtained by the processing can be superimposed on the second processing layer to obtain a processing result; the method includes:

[0010] Acquire first data, where the first data is represented by a plurality of images and category information corresponding to each image;

[0011] Training the convolutional neural network based on the first data;

[0012] The convolutional neural network trained based on the first data is used to identify the image to be identified, so as to obtain category information corresponding to the image to be identified.

[0013] In some embodiments, the sub-convolutional neural network layer includes a mixed convolutional layer; wherein the mixed convolutional layer includes: a sub-input layer, a sub-processing layer and five parallel sub-convolutional branches; the output end of the sub-input layer is respectively connected to the input end of the five parallel sub-convolutional branches, and the output ends of the five parallel sub-convolutional branches are respectively connected to the input end of the sub-processing layer; the data input into the mixed convolutional layer through the sub-input layer can be processed by the five parallel sub-convolutional branches respectively, and the processing results of the five sub-convolutional branches can be superimposed on the sub-processing layer to obtain the mixed convolutional layer processing result.

[0014] In some embodiments, the five parallel sub-convolution branches include: a first sub-convolution branch, a second sub-convolution branch, a third sub-convolution branch, a fourth sub-convolution branch and a fifth sub-convolution branch; wherein;

[0015] The first subconvolution branch includes a first convolution kernel and two second convolution kernels connected in sequence;

[0016] The second sub-convolution branch includes a first convolution kernel and two extended convolution kernels connected in sequence;

[0017] The third sub-convolution branch includes a first convolution kernel;

[0018] The fourth sub-convolution branch includes sequentially connecting an average pool and a first convolution kernel;

[0019] The fifth sub-convolution branch includes a first convolution kernel and two parallel second convolution kernels connected in sequence.

[0020] In some embodiments, the first processing layer includes two convolutional layers, a pooling layer, and a convolutional layer connected in sequence;

[0021] The second processing layer includes a convolution layer, a pooling layer and two convolution layers connected in sequence.

[0022] In some embodiments, the training of the convolutional neural network based on the first data includes:

[0023] constructing training data according to the first data;

[0024] Preprocess the constructed training data to obtain target training data;

[0025] The convolutional neural network is trained based on the target training data.

[0026] In some embodiments, the preprocessing of the constructed training data to obtain target training data includes:

[0027] The size of the image is compressed to a preset size.

[0028] The embodiment of the present application provides an image classification device based on a convolutional neural network, wherein the convolutional neural network includes a first processing layer, two or more parallel sub-convolutional neural network layers and a second processing layer; the output end of the first processing layer is respectively connected to the input end of the two or more parallel sub-convolutional neural network layers, and the output end of the two or more parallel sub-convolutional neural network layers is respectively connected to the input end of the second processing layer; after the data input into the first processing layer is processed by the first processing layer, it can be processed by the two or more parallel sub-convolutional neural network layers respectively, and the two or more sub-processing results obtained by the processing can be superimposed on the second processing layer to obtain a processing result; the device includes:

[0029] A data processing unit, configured to obtain first data, wherein the first data is represented by a plurality of images and category information corresponding to each image;

[0030] A training unit, configured to train the convolutional neural network based on the first data;

[0031] The recognition unit is used to recognize the image to be recognized by using the convolutional neural network trained based on the first data to obtain category information corresponding to the image to be recognized.

[0032] An embodiment of the present application provides a convolutional neural network, which includes: an input layer, an output layer, and two or more parallel sub-convolutional neural network layers; the output end of the input layer is respectively connected to the input ends of the two or more parallel sub-convolutional neural network layers; the input end of the output layer is respectively connected to the output ends of the two or more parallel sub-convolutional neural network layers; the data input to the input layer can be processed by the two or more parallel sub-convolutional neural network layers respectively, and the two or more sub-processing results obtained by processing can be superimposed on the input layer to obtain a processing result.

[0033] An embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the computer program is executed by the processor, the processor executes the steps of any image classification method based on a convolutional neural network in the above-mentioned embodiments.

[0034] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of any image classification method based on a convolutional neural network in the above-mentioned embodiments are implemented.

[0035] The method for gesture recognition in an embodiment of the present application obtains first data, wherein the first data is represented by multiple images and category information corresponding to each image; the convolutional neural network is trained based on the first data; the image to be recognized is recognized using the convolutional neural network trained based on the first data to obtain category information corresponding to the image to be recognized; the convolutional neural network includes a first processing layer, two or more parallel sub-convolutional neural network layers and a second processing layer connected in sequence; the present application proposes a new convolutional neural network, which can increase the number of feature value extractions through two or more parallel sub-convolutional neural network layers, reduce the amount of calculation in the data processing process, and improve the efficiency of data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings illustrate generally, by way of example and not limitation, various embodiments discussed herein.

[0037] Figure 1 This is a flow chart of an image classification method based on a convolutional neural network according to an embodiment of the present application;

[0038] Figure 2 This is a schematic diagram of the structure of a convolutional neural network in an embodiment of the present application;

[0039] Figure 3 This is a schematic diagram of the structure of a convolutional neural network in some embodiments of the present application;

[0040] Figure 4 Schematic diagram of the structure of the hybrid convolutional layer in the embodiment of the present application;

[0041] Figure 5 A schematic diagram of a system flow of a gesture recognition method according to some embodiments of the present application;

[0042] Figure 6 This is a schematic diagram of the effect after gesture image processing in some embodiments of the present application.

[0043] Figure 7 A schematic diagram of a system flow of an image classification method according to some embodiments of the present application;

[0044] Figure 8 This is a schematic diagram of the measured results of some embodiments of the present application;

[0045] Fig. 9 This is a schematic diagram of the measured results of some embodiments of the present application;

[0046] Fig.10 This is a schematic diagram of the structure of an image classification device based on a convolutional neural network according to an embodiment of the present application;

[0047] Fig.11 A schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0048] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.

[0049] In the embodiments of the present application, it should be noted that, unless otherwise specified and limited, the term "connection" should be understood in a broad sense. For example, it can be an electrical connection or a connection between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.

[0050] It should be noted that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects, and do not represent a specific order for the objects. It is understandable that the specific order or sequence of "first\second\third" can be interchanged where permitted. It should be understood that the objects distinguished by "first\second\third" can be interchanged where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.

[0051] Figure 1 This is a flow chart of an image classification method based on a convolutional neural network according to an embodiment of the present application. Figure 1 As shown, the image classification method based on convolutional neural network in the embodiment of the present application includes:

[0052] Step 101: Acquire first data, where the first data is represented by a plurality of images and category information corresponding to each image.

[0053] In some embodiments, the multiple images represented by the first data may include images of gestures, and correspondingly, the category information corresponding to each image may include the meaning corresponding to the gesture; for example: the gesture image of extending one finger can correspond to category information "1", the gesture image of extending two fingers can correspond to category information "2", etc.; this is just an example of the images represented by the first data and the corresponding category information, and is not a specific limitation on the embodiments of the present application; in actual applications, the category information corresponding to the specific image can be set according to user needs.

[0054] Step 102: training a convolutional neural network based on the first data.

[0055] In some embodiments, training a convolutional neural network based on the first data includes:

[0056] constructing training data based on the first data;

[0057] Preprocess the constructed training data to obtain target training data;

[0058] Train the convolutional neural network based on the target training data.

[0059] In some embodiments, preprocessing is performed on the constructed training data to obtain target training data, including: compressing the size of the image to a preset size.

[0060] In some embodiments, constructing training data based on the first data includes: processing multiple images represented by the first data to obtain a training set, and processing category information corresponding to the images to obtain labels corresponding to the images in the training set.

[0061] In some embodiments, preprocessing the constructed training data to obtain target training data also includes: grouping the training sets and corresponding labels in the training data according to a preset batch size of samples selected for one training, to obtain multiple batches of training sets and corresponding labels with a sample number of batch size.

[0062] The convolutional neural network is trained based on the target training data, including: inputting a plurality of batches of samples of batch size consisting of training sets and corresponding labels into the convolutional neural network in sequence for training.

[0063] Combining the training set and labels into a batch can improve training efficiency, make good use of the GPU, and increase training speed. Secondly, batch_size combined with the gradient descent algorithm can make the trained model more accurate. It avoids inputting all data into the network for training at one time, which causes huge differences in gradient values ​​in all directions caused by back propagation of calculated gradients. It improves global learning efficiency and reduces memory resource usage for a single training.

[0064] Step 103: Use the convolutional neural network trained based on the first data to identify the image to be identified, and obtain category information corresponding to the image to be identified.

[0065] It can be seen that in the embodiment of the present application, first data is obtained, and the first data is represented by multiple images and category information corresponding to each image; the convolutional neural network is trained based on the first data; the image to be identified is identified using the convolutional neural network trained based on the first data to obtain category information corresponding to the image to be identified; combined with the preprocessing of the training data, the target training data is obtained, which can improve the training efficiency of the convolutional neural network and make the trained convolutional neural network more accurate in distinguishing the category information corresponding to the image to be identified.

[0066] In order to solve the problem that the accuracy rate increases slowly or even decreases instead of increases due to the increasing depth of the convolutional neural network structure, this application studies the development direction of the convolutional neural network structure and designs a parallel convolutional neural network model. Unlike widening the network width by increasing the number of output channels of the convolutional layer, the data is input into independent convolutional neural networks for feature learning respectively, and then the obtained feature maps are fused. The basic module of the parallel neural network is a convolution model with a topological structure.

[0067] Figure 2 Schematic diagram of the structure of the convolutional neural network in the embodiment of the present application, such as Figure 2 As shown, in the embodiment of the present application, the convolutional neural network includes: a first processing layer 21, two or more parallel sub-convolutional neural network layers 22 and a second processing layer 23; wherein,

[0068] The output end of the first processing layer 21 is respectively connected to the input end of two or more parallel sub-convolutional neural network layers 22, and the output ends of the two or more parallel sub-convolutional neural network layers 22 are respectively connected to the input end of the second processing layer 23; after the data input into the first processing layer 21 is processed by the first processing layer 21, it can be processed by two or more parallel sub-convolutional neural network layers 22 respectively, and the two or more sub-processing results obtained by processing can be superimposed in the second processing layer 23 to obtain a processing result.

[0069] In the embodiment of the present application, the superposition of multiple processing results includes: superimposing multiple processing results in a third dimension. Taking the processing of image data as an example, the feature map obtained after the image data is processed by the convolutional neural network has an additional dimension relative to the original image data. In the present application, multiple parallel processing results are superimposed on this additional dimension; here, the third dimension is the dimension added to the feature map obtained by processing the image data by the convolutional neural network, that is, the dimension corresponding to the number of channels. Correspondingly, the size of the feature map can be expressed as the product of the width, height and number of channels.

[0070] In an embodiment of the present application, the feature map includes a processing result containing features obtained after the image data is processed by a convolutional neural network.

[0071] Figure 3 Schematic diagram of the structure of a convolutional neural network in some embodiments of the present application. In some embodiments, specifically, Figure 3 As shown, the first processing layer 21 includes two convolutional layers Conv, a pooling layer maxpool and a Conv connected in sequence. After each batch of data is input, it passes through two convolutional layers, a pooling layer and a convolutional layer. The main function is to extract features preliminarily before passing through the two branches and increase the number of channels of the feature map.

[0072] In some embodiments, the first processing layer 21 can increase the number of channels of the feature map from 3 to 64 to improve

[0073] In some embodiments of the present application, Figure 3 As shown in the figure, the convolution layer Conv in the convolutional neural network uses a 3×3 convolution kernel, namely Conv3×3; compared with the large convolution kernels of 5×5, 7×7 and 11×11, the 3×3 convolution kernel significantly reduces the number of parameters, thereby improving the computing performance of the convolutional neural network.

[0074] In some embodiments, the second processing layer 23 includes a Conv, a maxpool and two Convs connected in sequence. The second processing layer 23 can superimpose the two or more sub-processing results obtained by processing the above two or more parallel sub-convolutional neural network layers 22 in the third dimension. The last two Convs of the second processing layer 23 replace the fully connected layer with too many parameters, and the kernel size is gradually reduced and the convolution with padding=VALID is filled, which just reduces the width and height of the feature map to 1x1, so that the calculation amount is smaller, the convergence is faster, and the number of parameters is less.

[0075] When deploying deep learning models on mobile devices, we must not only consider the accuracy of the model, but also pay attention to the number of model parameters.

[0076] In some embodiments, the sub-convolutional neural network layer 22 includes a mixed convolutional layer Conv-mixed, two Convs, one Conv-mixed, two Convs, one Conv-mixed and one Conv connected in sequence.

[0077] Among them, the mixed convolution layer Conv-mixed is a new convolution layer proposed in this application. Specifically, Figure 4 Schematic diagram of the structure of the hybrid convolutional layer in the embodiment of the present application. Figure 4 As shown, the hybrid convolution layer includes: a sub-input layer 41 (Previous layer), a sub-processing layer 42 (Concatenation) and five parallel sub-convolution branches 401 to 405; the output end of the sub-input layer 41 is respectively connected to the input ends of the five parallel sub-convolution branches 401 to 405, and the output ends of the five parallel sub-convolution branches 401 to 405 are respectively connected to the input end of the sub-processing layer 42; the data input into the hybrid convolution layer through the sub-input layer 42 can be processed by the five parallel sub-convolution branches 401 to 405 respectively, and the processing results of the five sub-convolution branches 401 to 405 can be superimposed on the sub-processing layer 42 to obtain the processing result of the hybrid convolution layer.

[0078] In the embodiment of the present application, the superposition of multiple processing results includes: superimposing multiple processing results in a third dimension. Taking the processing of image data as an example, the feature map obtained after the image data is processed by the convolutional neural network has an additional dimension relative to the original image data. In the present application, multiple parallel processing results are superimposed on this additional dimension; here, the third dimension is the dimension added to the feature map obtained by processing the image data by the convolutional neural network, that is, the dimension corresponding to the number of channels. Correspondingly, the size of the feature map can be expressed as the product of the width, height and number of channels.

[0079] In some embodiments, Figure 4 As shown, the five parallel sub-convolution branches 401 to 405 include: a first sub-convolution branch 401, a second sub-convolution branch 402, a third sub-convolution branch 403, a fourth sub-convolution branch 404 and a fifth sub-convolution branch 405; wherein;

[0080] The first subconvolution branch 401 includes a first convolution kernel and two second convolution kernels connected in sequence;

[0081] The second sub-convolution branch 402 includes a first convolution kernel and two dilated convolution kernels dilatedconv connected in sequence;

[0082] The third sub-convolution branch 403 includes a first convolution kernel;

[0083] The fourth sub-convolution branch 404 includes an average pool Avg_pool and a first convolution kernel connected in sequence;

[0084] The fifth sub-convolution branch 405 includes a first convolution kernel and two parallel second convolution kernels connected in sequence.

[0085] In some embodiments, the first convolution kernel includes conv(k:1×1, s:1), the second convolution kernel includes conv(k:3×3, s:1), the extended convolution kernel includes dilated conv(k:3×3, s:2, r:2), and the average pool includes Avg_pool(k:3×3, s:1).

[0086] It should be noted that Figure 4 The numbers C0 to C10 in some specific embodiments of the present application are used to facilitate the distinction of the various functional modules in each sub-convolution branch in the mixed convolution layer Conv-mixed. They are only understood as numbers, rather than limiting the names of the functional modules in each sub-convolution branch.

[0087] When deploying deep learning models on mobile devices, we must not only consider the accuracy of the model, but also pay attention to the number of model parameters. Figure 5Parameter diagram of convolutional layers in convolutional neural networks in some embodiments of the present application, where the specific parameters of each convolutional layer are Figure 5 Each branch uses a small convolution kernel to increase the number of channels. The number of parameters and calculations of this branch topology is only 1 / 3 of the traditional linear structure, which can further improve the computing performance of the convolutional neural network.

[0088] In some embodiments, in order to study the performance of convolutional neural networks with independent branch structures and traditional neural networks, the present application embodiment proposes the following Figure 3 The convolutional neural network structure shown in the figure, after each batch of data is input, it passes through two convolutional layers, a pooling layer and a convolutional layer. Its main function is to extract features before passing through the two branches, and increase the number of channels of the feature map from 3 to 64, and then pass through the parallel structure composed of two branches. After the branch is completed, two feature maps are superimposed in the third dimension. The last two convolutions mainly replace the fully connected layer with too many parameters. The kernel size gradually decreases, and the convolution with padding = VALID just reduces the width and height of the feature map to 1x1, which reduces the amount of calculation, converges faster, and requires fewer parameters. The parameters of the convolutional neural network are as follows: Figure 5 As shown, Figure 5 The detailed parameters of conv, maxpool and Conv-mixed are introduced in the text, such as the third column is the product of the width, height and number of channels of the output size of each layer; the fourth column Filter size / Stride represents the kernel size and stride of ordinary convolution and maximum pooling; the basic parameters of Conv-mixed are in the fifth column Feature maps (Conv-mixed); the x2 in the last column represents a parallel structure, and the parameters in the parallel structure are consistent.

[0089] Assuming that the input of the convolutional neural network is x, the various convolution, regularization, pooling and other operations in the middle can be regarded as a complex function F(x). When there is only a single branch structure, the expected result can be described as:

[0090] H(x)=F(x)

[0091] As the number of branches increases, the expected results can be described as:

[0092]

[0093] Where i represents the number of branches in the parallel structure. Although F(x) can be fitted as However, experiments show that it takes much longer and more iterations to train the model to achieve the same effect as the former, and the network converges more slowly.

[0094] The number of parameters of the convolutional neural network with the above parallel structure is calculated as follows:

[0095]

[0096] where k l and k d Indicates the width and height of the current convolution kernel, N i-1 Represents the input of the current convolutional layer, that is, the output feature map depth of the previous layer, b represents the bias, n represents the depth of the branch in the parallel structure, and N represents the number of branches in the parallel structure.

[0097] In order to illustrate the effect of the image classification method of the convolutional neural network in the embodiment of the present application, the present application provides some examples of gesture image classification. Specifically,

[0098] The dataset has been preprocessed in the early stage, including data enhancement and compression. After being combined into batches, it is trained in a deep convolutional neural network. When compressing images, you cannot simply directly resize the image, which will cause the image to be distorted, seriously reduce the quality of the training set, and lead to extremely poor test results. Here, the area interpolation method is used to resize the image to 64×64×3. The effect of the processed gesture image is as follows: Figure 6 shown.

[0099] The whole system process is as follows Figure 7 Although Tensorflow supports mobile training, the computing power of mobile terminals is limited and the training time is too long, so the training process can be performed on the server. Figure 3 When the convolutional neural network shown in the figure is used, batch_size is an important parameter, which indicates the number of samples required for one iteration of the neural network. If batch_size is not used, but all data is input into the network for training at one time to calculate the gradient for back propagation, the gradient values ​​obtained will vary greatly in all directions, making it difficult to use a good global learning rate in advance. In addition, the amount of data input at one time is huge, and the memory resource requirements occupied are also difficult to meet. Combining the training set and the label into a batch can improve the training efficiency, make good use of the GPU, and increase the training speed. In addition, batch_size combined with the gradient descent algorithm can train the model with higher accuracy.

[0100] After the model training is completed, the model is persisted and saved as a .pb file, and then the model file and label file are placed in the corresponding location of the project. When the program is running, open the mobile camera to take pictures, use the area interpolation method to compress and combine them into batches, use the library function to call the model and input the batch data into the model, use the Fetch() function to get the output result, and display it on the front-end interface. When calling the trained model, use the corresponding library function to call it directly. The measured results are as follows: Figure 8 , Fig. 9 shown.

[0101] Lightweight deep convolutional neural networks have the characteristics of small number of parameters and small amount of calculation, which is only 1 / 3 of the traditional deep convolutional neural network structure, and can achieve better results. Compared with the existing convolutional neural network algorithms, it has stronger convergence.

[0102] The current convolutional neural network structure has a large number of parameters and high computational complexity, which is not conducive to deployment on mobile platforms. This application proposes a lightweight deep convolutional neural network gesture recognition algorithm with a small number of parameters, small computational complexity, strong convergence, high portability, and easy training.

[0103] Fig.10 Schematic diagram of the structure of an image classification device based on a convolutional neural network according to an embodiment of the present application. Fig.10 As shown, the device is based on the convolutional neural network described in the embodiment of the present application, and the convolutional neural network of the aforementioned embodiment of the present application is not described here; the device includes: a data processing unit 31, a training unit 32 and a recognition unit 33. Among them,

[0104] The data processing unit 31 is used to obtain first data, where the first data is represented by a plurality of images and category information corresponding to each image.

[0105] The training unit 32 is used to train the convolutional neural network based on the first data.

[0106] In some embodiments, the training of the convolutional neural network based on the first data includes:

[0107] constructing training data according to the first data;

[0108] Preprocess the constructed training data to obtain target training data;

[0109] The convolutional neural network is trained based on the target training data.

[0110] In some embodiments, the preprocessing of the constructed training data to obtain target training data includes: compressing the size of the image to a preset size.

[0111] The recognition unit 33 is used to recognize the image to be recognized by using the convolutional neural network trained based on the first data to obtain category information corresponding to the image to be recognized.

[0112] The present application also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, it is used to implement at least the steps of the image classification method based on a convolutional neural network in the above embodiment. The computer-readable storage medium can be a memory. The memory can be Fig.11 The memory 82 is shown.

[0113] An embodiment of the present application also provides an electronic device. Fig.11 Schematic diagram of the hardware structure of the electronic device of the embodiment of the present application, such as Fig.11 As shown, the terminal includes: a communication component 83 for data transmission, at least one processor 81, and a memory 82 for storing a computer program that can be run on the processor 81. The various components in the terminal are coupled together through a bus system 84. It can be understood that the bus system 84 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 84 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig.11 Various buses are labeled as bus system 84 .

[0114] When the processor 81 executes the computer program, at least the steps of the image classification method based on convolutional neural network in the aforementioned embodiment are executed.

[0115] It can be understood that the memory 82 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAM bus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 82 described in the embodiments of the present application is intended to include, but is not limited to, these and any other suitable types of memory.

[0116] The method disclosed in the above embodiment of the present application can be applied to the processor 81, or implemented by the processor 81. The processor 81 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 81 or the instruction in the form of software. The above processor 81 can be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 81 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 82, and the processor 81 reads the information in the memory 82 and completes the steps of the above method in combination with its hardware.

[0117] In an exemplary embodiment, the relevant device or classification system can be implemented by one or more application-specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), FPGA, general-purpose processor, controller, MCU, microprocessor, or other electronic components to execute the aforementioned classification model training method and / or classification method.

[0118] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0119] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0120] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0121] A person skilled in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, etc. Various media that can store program codes.

[0122] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0123] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0124] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0125] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0126] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. An image classification method based on convolutional neural network, characterized in that: The convolutional neural network comprises: a first processing layer, two or more parallel sub-convolutional neural network layers and a second processing layer; the output end of the first processing layer is respectively connected to the input end of the two or more parallel sub-convolutional neural network layers, and the output end of the two or more parallel sub-convolutional neural network layers is respectively connected to the input end of the second processing layer; after the data input into the first processing layer is processed by the first processing layer, it can be processed by the two or more parallel sub-convolutional neural network layers respectively, and the two or more sub-processing results obtained by the processing can be superimposed on the second processing layer to obtain a processing result; the method comprises: Acquire first data, where the first data is represented by a plurality of images and category information corresponding to each image; Training the convolutional neural network based on the first data; The convolutional neural network trained based on the first data is used to identify the image to be identified, so as to obtain category information corresponding to the image to be identified.

2. The method according to claim 1, characterized in that The sub-convolutional neural network layer includes a mixed convolutional layer; wherein the mixed convolutional layer includes: a sub-input layer, a sub-processing layer and five parallel sub-convolutional branches; the output end of the sub-input layer is respectively connected to the input end of the five parallel sub-convolutional branches, and the output ends of the five parallel sub-convolutional branches are respectively connected to the input end of the sub-processing layer; the data input into the mixed convolutional layer through the sub-input layer can be processed by the five parallel sub-convolutional branches respectively, and the processing results of the five sub-convolutional branches can be superimposed on the sub-processing layer to obtain the mixed convolutional layer processing result.

3. The method according to claim 2, characterized in that The five parallel sub-convolution branches include: a first sub-convolution branch, a second sub-convolution branch, a third sub-convolution branch, a fourth sub-convolution branch and a fifth sub-convolution branch; wherein; The first subconvolution branch includes a first convolution kernel and two second convolution kernels connected in sequence; The second sub-convolution branch includes a first convolution kernel and two extended convolution kernels connected in sequence; The third sub-convolution branch includes a first convolution kernel; The fourth sub-convolution branch includes sequentially connecting an average pool and a first convolution kernel; The fifth sub-convolution branch includes a first convolution kernel and two parallel second convolution kernels connected in sequence.

4. The method according to claim 2, characterized in that: The first processing layer includes two convolutional layers, a pooling layer and a convolutional layer connected in sequence; The second processing layer includes a convolution layer, a pooling layer and two convolution layers connected in sequence.

5. The method according to claim 1, characterized in that The training of the convolutional neural network based on the first data includes: constructing training data according to the first data; Preprocess the constructed training data to obtain target training data; The convolutional neural network is trained based on the target training data.

6. The method according to claim 5, characterized in that The preprocessing of the constructed training data to obtain target training data includes: The size of the image is compressed to a preset size.

7. An image classification device based on convolutional neural network, characterized in that: The convolutional neural network includes a first processing layer, two or more parallel sub-convolutional neural network layers and a second processing layer; the output end of the first processing layer is respectively connected to the input end of the two or more parallel sub-convolutional neural network layers, and the output ends of the two or more parallel sub-convolutional neural network layers are respectively connected to the input end of the second processing layer; after the data input into the first processing layer is processed by the first processing layer, it can be processed by the two or more parallel sub-convolutional neural network layers respectively, and the two or more sub-processing results obtained by the processing can be superimposed on the second processing layer to obtain a processing result; the device includes: A data processing unit, configured to obtain first data, wherein the first data is represented by a plurality of images and category information corresponding to each image; A training unit, configured to train the convolutional neural network based on the first data; The recognition unit is used to recognize the image to be recognized by using the convolutional neural network trained based on the first data to obtain category information corresponding to the image to be recognized.

8. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Multiscale feature convolutional neural network-based image scene classification method

    CN108491856A

  • Data processing method, device, computer, and storage medium

    CN109409504A