Image classification system based on deep convolutional neural network and classification method thereof
By designing an image classification system based on deep convolutional neural networks with multiple modules, the problems of large computing, complex processing and poor prediction accuracy in image classification applications are solved, and efficient and accurate image classification and deployment adjustment are achieved.
Patent Information
- Application Number
- CN202510171538.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Image classification application based on deep convolutional neural networks has large computational volume, complex processing, poor data prediction accuracy, and difficult to adjust data after deployment in tasks such as object detection, image segmentation and target tracking.
An image classification system based on deep convolutional neural network is designed, including data preparation module, model generation module, model training module, model evaluation module and model optimization and deployment module. The system achieves efficient and accurate image classification through data preprocessing, model architecture selection, network structure definition, model training and evaluation, as well as model optimization and deployment.
It achieves small calculation volume, simple processing, high data prediction accuracy, and can continue to adjust data after deployment, solving the problems of large calculation volume, complex processing and poor prediction accuracy in the original technology.
Smart Images

Figure CN120047745A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image classification, and specifically to an image classification system based on a deep convolutional neural network and its classification method. Background Art
[0002] Image classification technology is a prerequisite for complex computer vision tasks. In particular, tasks such as object detection, image segmentation, and object tracking all require image classification as a pre-task. Currently, image classification generally uses a trained image classification neural network. In recent years, deep convolutional neural networks have achieved great success in many computer vision tasks. Especially in the image classification task, numerous neural networks with good performance have been designed. Among them, the representative Residual Network (ResNet) has proven the effectiveness of very deep neural networks. Therefore, a large number of extension methods have been continuously improved based on it, such as Xception, WideResNet, PyramidNet, and ResNeXt.
[0003] Currently, when applying image classification based on a deep convolutional neural network to image classification such as object detection, image segmentation, and object tracking, there are problems such as large computational complexity, complex processing, poor accuracy of data prediction, and difficulty in continuously adjusting data after deployment. Therefore, we propose an image classification system based on a deep convolutional neural network and its classification method. Summary of the Invention
[0004] The purpose of the present invention is to provide an image classification system based on a deep convolutional neural network and its classification method, which has the advantages of small computational complexity, simple processing, high accuracy of data prediction, and the ability to continue adjusting data after deployment, and solves the problems of large computational complexity, complex processing, poor accuracy of data prediction, and difficulty in continuously adjusting data after deployment when applying image classification based on a deep convolutional neural network to image classification such as object detection, image segmentation, and object tracking.
[0005] To achieve the above purpose, the present invention provides the following technical solution: An image classification system based on a deep convolutional neural network, including a data preparation module, a model generation module, a model training module, a model evaluation module, and a model optimization and deployment module. The output end of the data preparation module is signal-connected to the input end of the model generation module, the output end of the model generation module is signal-connected to the input end of the model training module, the output end of the model training module is signal-connected to the input end of the model evaluation module, and the output end of the model evaluation module is signal-connected to the input end of the model optimization and deployment module; When inputting new data and performing classification prediction of new images, when the result prediction is correct, add the correct result to the model generation module and repeat the training again; When the result prediction is incorrect, manually calibrate the incorrect result and add the calibrated correct result to the data preparation module, and repeat the training again.
[0006] Preferably, the data preparation module includes: Data collection unit: used to collect a large amount of image data to ensure the diversity and representativeness of the data set; Data annotation unit: used to classify and annotate images to ensure that each image has a corresponding class label; Data preprocessing unit: used to convert the original image data into a form suitable for subsequent processing and analysis, to improve the efficiency and accuracy of image processing, reduce the amount of calculation and improve the calculation efficiency.
[0007] Preferably, the model generation module includes: Model architecture selection unit: According to the usage needs, adopt pre-trained models of VGG, ResNet, and Inception, which are used as the basis of the model architecture and perform transfer learning; Network structure definition unit: used to effectively extract image features and realize tasks such as image classification, detection, and segmentation; Hardware preparation unit: used to improve computing performance and system stability, and also ensure the scalability and compatibility of the system through high-performance hardware configuration, so as to achieve more efficient and reliable image classification tasks in actual applications.
[0008] Preferably, the model training module includes: Dataset division unit: used to divide the dataset into training set, validation set, and test set; Hyperparameter setting unit: used to manually set the parameters required during network training according to the actual problem and dataset; Model compilation unit: used to select loss function, optimizer, and evaluation metrics.
[0009] Preferably, the model evaluation module includes: Validation set evaluation unit: used to monitor the performance of the model using the validation set during training to adjust hyperparameters; Test set evaluation unit: used to evaluate the final performance of the model on an independent test set to calculate metrics such as accuracy, precision, recall, and F1 score; Evaluation metrics: used to evaluate the performance of the model; Confusion matrix: used to analyze the performance of the model on different classes; Visualization: used to plot the loss and accuracy curves during training.
[0010] Preferably, the model optimization and deployment module includes: Model Optimization Unit: Optimizes the model using techniques such as pruning, quantization, and knowledge distillation to reduce inference time; Storage Unit: Used to save the trained model as a file; Deployment Environment Unit: Used to select a suitable deployment environment; Loading Unit: Used to load the model in the deployment environment; Integration Unit: Used to integrate the model into the application to achieve the function of real-time image classification.
[0011] Preferably, the data preprocessing unit includes: Scaling: Used to resize all images to the same size; Normalization: Used to map the input data to a standard normal distribution with a mean of 0 and a variance of 1, facilitating the learning of subsequent layers, making the data distribution more stable, and benefiting the stability and generalization ability of network training; Data Augmentation: According to the data requirements, rotates, scales, crops, and flips the data to increase data diversity, used to improve the generalization ability of the model.
[0012] Preferably, the network structure definition unit includes: Convolutional Layer: Used to extract the features of the input data, and by sliding the convolutional kernel on the input data, calculates the dot product of the convolutional kernel and the local area of the input data to generate a feature map; Pooling Layer: Used to reduce the spatial dimension of the feature map, reduce the number of parameters, and improve the generalization ability of the model; Fully Connected Layer: Used to convert the feature map into the final output result and map the features to class labels; Activation Function: Used to introduce non-linearity and enhance the expressive ability of the model; Regularization: Used to prevent overfitting of the convolutional neural network.
[0013] Preferably, the hyperparameter setting unit and the model compilation unit include: Learning Rate: Used to determine the step size of the model when updating the weights each time; Batch Size: Used to determine the number of samples used in each training; Number of Iterations: Used to determine the total number of training epochs.
[0014] An image classification method based on a deep convolutional neural network, the specific steps include: S1: Establish stackable basic blocks for stacking to form neural architectures of different depths, and classify each basic block data according to the image type; S2: Construct a custom deep convolutional neural network; S3: Set predetermined training parameters, use the data for training, and use the test data for testing. The specific implementation method includes the following steps: a. After obtaining the training data, preprocess the data; b. Determine the parameter data for the predetermined training of the network according to the preprocessed data; c. Use the parameter data for the predetermined training to train the network to obtain a detailed classification model; d. Test the obtained detailed classification model, quickly distinguish the differences in various image details including image color, image position, image size, image content, and image action, classify the differences in details, determine the optimization method, and obtain a detailed classification equation; S4: Input new data for prediction.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention can generate corresponding models after data preparation, establish a stackable neural architecture for stacking to form different depths, and classify each basic block data according to image classification types with classification labels. Moreover, the data preprocessing after data annotation can convert the original image data into a form suitable for subsequent processing and analysis, so as to improve the efficiency and accuracy of image processing, reduce the amount of calculation and improve the calculation efficiency. The use of the model evaluation module and the model optimization and deployment module enables the system to independently evaluate the final performance of the model on the test set. And the use of techniques such as pruning, quantization, and knowledge distillation to optimize the model can reduce the inference time, ensure the accuracy of the system during use, and after deployment, when inputting new data for classification prediction, if there is a prediction error, it can provide image data for the data preparation unit synchronously through manual calibration, which is beneficial to the learning and growth of the system model.
[0016] 2. The present invention can grow rapidly during use, is convenient and flexible to use, has high accuracy, can quickly find the differences on similar images in details, and can quickly distinguish and classify images according to different differences, which can avoid the situation that similar images cannot be quickly distinguished or there are classification errors for similar images, and can quickly query for the user where the differences in the image are, which is convenient for the user to use. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the system flow of the present invention; Figure 2 It is a schematic diagram of the structure of the data preparation module of the present invention; Figure 3 It is a schematic diagram of the structure of the model generation module of the present invention; Figure 4 It is a schematic diagram of the structure of the model training module of the present invention; Figure 5 Schematic diagram of the model evaluation module structure of the present invention; Figure 6 Schematic diagram of the model optimization and deployment module structure of the present invention. Specific implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] Please refer to Figures 1-6 As shown, the present invention provides a technical solution: an image classification system based on a deep convolutional neural network, including a data preparation module, a model generation module, a model training module, a model evaluation module, and a model optimization and deployment module. The output end of the data preparation module is signal-connected to the input end of the model generation module, the output end of the model generation module is signal-connected to the input end of the model training module, the output end of the model training module is signal-connected to the input end of the model evaluation module, and the output end of the model evaluation module is signal-connected to the input end of the model optimization and deployment module; When inputting new data for classification prediction of new images, when the result prediction is correct, add the correct result to the model generation module and repeat the training again; When the result prediction is incorrect, manually calibrate the incorrect result and add the calibrated correct result to the data preparation module and repeat the training again.
[0020] This technical solution: can generate a corresponding model after data preparation, establish a stackable neural architecture for stacking to form different depths, and classify each basic block data according to image classification, and the data preprocessing after data annotation can convert the original image data into a form suitable for subsequent processing and analysis to improve the efficiency and accuracy of image processing, reduce the amount of calculation and improve the calculation efficiency. The use of the model evaluation module and the model optimization and deployment module enables this system to independently evaluate the final performance of the model on the test set, and uses techniques such as pruning, quantization, and knowledge distillation to optimize the model, which can reduce the inference time, ensure the accuracy when the system is used, and after deployment, when inputting new data for classification prediction, if the prediction is incorrect, it can synchronously provide image data for the data preparation unit through manual calibration, which is beneficial to the learning and growth of the model of this system.
[0021] Specifically, the data preparation module includes: Data collection unit: used to collect a large amount of image data to ensure the diversity and representativeness of the dataset; Data annotation unit: used to classify and annotate images to ensure that each image has a corresponding class label; Data preprocessing unit: used to convert the original image data into a form suitable for subsequent processing and analysis, so as to improve the efficiency and accuracy of image processing, reduce the computational amount and improve the computational efficiency.
[0022] Specifically, the model generation module includes: Model architecture selection unit: adopt pre-trained models of VGG, ResNet, Inception according to usage needs, which are used as the basis of the model architecture and perform transfer learning; Taking ResNet as an example, based on the existing ResNet technology, train a ResNet model that conforms to the scenario according to the recognition method.
[0023] ResNet solves the gradient vanishing problem of deep networks by connecting each convolutional block with the output of the previous block. This connection method is called a residual connection, which enables the network to learn the feature expression ability of more layers. The core concept of ResNet is the residual block, which consists of multiple convolutional layers and Batch Normalization layers.
[0024] The core algorithm principle of ResNet is to solve the gradient vanishing problem of deep networks through residual connections. In ResNet, each convolutional block has a residual connection, which adds the output of the current block to the output of the previous block and then activates it through an activation function. This connection method enables the network to learn the feature expression ability of more layers.
[0025] The specific operation steps are as follows: 1. The input image learns initial features through a convolutional layer and a Batch Normalization layer.
[0026] 2. These features will be passed to the first residual block.
[0027] 3. In each residual block, the input features are connected with the output of the previous block through a residual connection.
[0028] 4. The result of the residual connection is activated through an activation function (such as ReLU).
[0029] 5. The activated features will be passed to the next convolutional block.
[0030] 6. This process will continue until the last convolutional block.
[0031] 7. The output of the last convolutional block is classified through a fully connected layer and a Softmax activation function.
[0032] Network structure definition unit: used to effectively extract image features and implement tasks such as image classification, detection, and segmentation; Hardware preparation unit: used to improve computing performance and system stability, and also ensure the scalability and compatibility of the system through high-performance hardware configurations, so as to achieve more efficient and reliable image classification tasks in practical applications.
[0033] It can be understood that using a GPU in the hardware preparation unit can accelerate the training process, and NVIDIA GPUs are recommended. In the model architecture selection unit, a new CNN architecture can also be designed separately.
[0034] Specifically, the model training module includes: Dataset division unit: used to divide the dataset into a training set, a validation set, and a test set; Hyperparameter setting unit: used to manually set parameters required during network training according to the actual problem and the dataset; Model compilation unit: used to select a loss function, an optimizer, and evaluation metrics.
[0035] It can be understood that in the model compilation unit, a loss function such as cross-entropy, an optimizer such as Adam, and evaluation metrics such as accuracy are selected.
[0036] Specifically, the model evaluation module includes: Validation set evaluation unit: used to monitor the performance of the model using the validation set during training to adjust hyperparameters; Test set evaluation unit: used to evaluate the final performance of the model on an independent test set to calculate metrics such as accuracy, precision, recall, and F1 score; Evaluation metrics: used to evaluate the performance of the model; Confusion matrix: used to analyze the performance of the model on different classes; Visualization: used to plot the loss and accuracy curves during training.
[0037] It can be understood that the test set evaluation unit is mainly used to evaluate the final performance and generalization ability of the model, and thus is the final test of the model's generalization ability to ensure that the model can work stably in actual scenarios.
[0038] Specifically, the model optimization and deployment module includes: Model optimization unit: uses techniques such as pruning, quantization, and knowledge distillation to optimize the model to reduce inference time; Storage unit: used to save the trained model as a file; Deployment environment unit: used to select a suitable deployment environment; Loading unit: used to load the model in the deployment environment; Integration unit: used to integrate the model into the application to achieve the classification function of real-time images.
[0039] Specifically, the data preprocessing unit includes: Scaling: used to resize all images to the same size; Normalization: used to map the input data to a standard normal distribution with a mean of 0 and a variance of 1, facilitating the learning of subsequent layers, making the data distribution more stable, and being conducive to the stability and generalization ability of network training; Data augmentation: according to the data requirements, rotation, scaling, cropping, and flipping are used to increase data diversity, which is used to improve the generalization ability of the model.
[0040] It can be understood that the main purpose of the data preprocessing unit is to improve the accuracy, reliability, and efficiency of data analysis and modeling. Specifically, the main purposes of data preprocessing include data cleaning, data transformation, data integration, data normalization, and data dimensionality reduction, thereby ensuring the quality of the data, including ensuring the accuracy, integrity, and consistency of the data.
[0041] Specifically, the network structure definition unit includes: Convolutional layer: used to extract the features of the input data, and by sliding the convolutional kernel on the input data, calculate the dot product of the convolutional kernel and the local area of the input data to generate a feature map; Pooling layer: used to reduce the spatial dimension of the feature map, reduce the number of parameters, and improve the generalization ability of the model; Fully connected layer: used to convert the feature map into the final output result and map the features to class labels; Activation function: used to introduce non-linearity and enhance the expressive ability of the model; Regularization: used to prevent overfitting of the convolutional neural network.
[0042] It is understandable that the commonly used pooling operations include Max Pooling and Average Pooling. Max Pooling retains the most important features by taking the maximum value in the local area, while Average Pooling smooths the features by calculating the average value in the local area. Regularization methods include L1 regularization, L2 regularization, Dropout, etc. L1 and L2 regularization limit the size of model parameters by adding regularization terms to the loss function. Dropout increases the generalization ability of the model by randomly discarding neurons in the network. Commonly used activation functions include ReLU, Sigmoid, Tanh, etc. The ReLU function is widely used in convolutional neural networks due to its advantages such as simple calculation and fast training speed.
[0043] Specifically, the hyperparameter setting unit and the model compilation unit include: Learning rate: used to determine the step size of the model when updating weights each time; Batch size: used to determine the number of samples used in each training; Number of iterations: used to determine the total number of training epochs.
[0044] It is understandable that if the learning rate is set too high, the model may not converge; if the learning rate is set too low, the model may converge too slowly. The choice of batch size affects the convergence speed and memory usage of the model.
[0045] An image classification method based on a deep convolutional neural network, the specific steps include: S1: Establish stackable basic blocks for stacking to form neural architectures of different depths, and assign classification labels to the data of each basic block according to image classification; S2: Construct a custom deep convolutional neural network; S3: Set predetermined training parameters, train using data, and test using test data. The specific implementation method includes the following steps: a. After obtaining the training data, preprocess the data; b. According to the preprocessed data, determine the parameter data for the predetermined training of the network; c. Use the parameter data for the predetermined training to train the network to obtain a detailed classification model; d. Test the obtained detailed classification model to quickly distinguish the differences in various image details including image color, image position, image size, image content, and image actions, classify the differences in details, determine the optimization method, and obtain a detailed classification equation; S4: Input new data for prediction.
[0046] This technical solution: can grow rapidly during use, is convenient and flexible to use, has high accuracy, can quickly find the differences on similar images in details, and classify the images quickly according to different differences, can avoid the situation that similar images cannot be quickly distinguished or the classification of similar images goes wrong, and can quickly query for the user the location where the differences in the image are, which facilitates the use of the user.
[0047] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. An image classification system based on a deep convolutional neural network, comprising a data preparation module, a model generation module, a model training module, a model evaluation module and a model optimization and deployment module, characterized in that: The output end of the data preparation module is signal-connected to the input end of the model generation module, the output end of the model generation module is signal-connected to the input end of the model training module, the output end of the model training module is signal-connected to the input end of the model evaluation module, and the output end of the model evaluation module is signal-connected to the input end of the model optimization and deployment module; When new data is input and classification prediction is performed on new images, if the prediction result is correct, the correct result is added to the model generation module and training is repeated again; When the result prediction is wrong, the wrong result is manually calibrated, and the calibrated correct result is added to the data preparation module, and the training is repeated again.
2. The image classification system based on deep convolutional neural network according to claim 1, characterized in that: The data preparation module includes: Data collection unit: used to collect a large amount of image data to ensure that the data set is diverse and representative; Data annotation unit: used to classify and annotate images to ensure that each image has a corresponding category label; Data preprocessing unit: used to convert raw image data into a form suitable for subsequent processing and analysis to improve the efficiency and accuracy of image processing, reduce the amount of calculation and improve calculation efficiency.
3. The image classification system based on deep convolutional neural network according to claim 1, characterized in that: The model generation module includes: Model architecture selection unit: Use VGG, ResNet, and Inception pre-trained models as needed as the basis for model architecture and perform transfer learning; Network structure definition unit: used to effectively extract image features and achieve image classification, detection and segmentation tasks; Hardware preparation unit: used to improve computing performance and system stability, and also ensure the scalability and compatibility of the system through high-performance hardware configuration, so as to achieve more efficient and reliable image classification tasks in practical applications.
4. The image classification system based on deep convolutional neural network according to claim 1, characterized in that: The model training module includes: Dataset division unit: used to divide the dataset into training set, validation set and test set; Hyperparameter setting unit: used to manually set parameters during network training according to actual problems and data sets; Compile model unit: used to select loss function, optimizer and evaluation metric.
5. The image classification system based on deep convolutional neural network according to claim 1, characterized in that: The model evaluation module includes: Validation set evaluation unit: used to monitor the performance of the model using the validation set during training to adjust the hyperparameters; Test set evaluation unit: used to evaluate the final performance of the model on an independent test set to calculate the accuracy, precision, recall, and F1 score indicators; Evaluation metrics: used to evaluate the performance of the model; Confusion matrix: used to analyze the performance of the model on different categories; Visualization: Used to plot loss and accuracy curves during training.
6. The image classification system based on deep convolutional neural network according to claim 1, characterized in that: The model optimization and deployment module includes: Model optimization unit: uses pruning, quantization, and knowledge distillation techniques to optimize the model to reduce inference time; Storage unit: used to save the trained model as a file; Deployment environment unit: used to select a suitable deployment environment; Loading unit: used to load the model in the deployment environment; Integration unit: used to integrate the model into the application to achieve real-time image classification capabilities.
7. The image classification system based on deep convolutional neural network according to claim 2, characterized in that: The data preprocessing unit comprises: Scale: used to resize all images to the same size; Normalization: used to map the input data to a standard normal distribution with a mean of 0 and a variance of 1, so as to facilitate the learning of subsequent layers and make the data distribution more stable, which is conducive to the stability and generalization ability of network training; Data enhancement: According to data needs, data diversity can be increased by rotating, scaling, cropping, and flipping to improve the generalization ability of the model.
8. The image classification system based on deep convolutional neural network according to claim 3, characterized in that: The network structure definition unit includes: Convolutional layer: used to extract the features of the input data, and calculate the dot product between the convolution kernel and the local area of the input data by sliding the convolution kernel on the input data to generate a feature map; Pooling layer: used to reduce the spatial dimension of the feature map, reduce the number of parameters, and improve the generalization ability of the model; Fully connected layer: used to convert feature maps into final output results and map features to category labels; Activation function: used to introduce nonlinearity and enhance the expressiveness of the model; Regularization: Used to prevent overfitting of convolutional neural networks.
9. The image classification system based on deep convolutional neural network according to claim 4, characterized in that: The hyperparameter setting unit and the model compilation unit include: Learning rate: used to determine the step size of the model each time the weight is updated; Batch size: used to determine the number of samples used in each training; Iterations: used to determine the total number of training rounds.
10. An image classification method based on deep convolutional neural network, characterized in that: The image classification method comprises an image classification system based on a deep convolutional neural network according to any one of claims 1 to 9, and the specific steps include: S1: Establish stackable basic blocks for stacking to form neural architectures of different depths, and classify and label each basic block data according to image classification; S2: Build a custom deep convolutional neural network; S3: Setting predetermined training parameters, using data for training, and using test data for testing, the specific implementation method includes the following steps: a. After obtaining the training data, preprocess the data; b. Determine the parameter data of the network scheduled training based on the preprocessed data; c. Train the network using the predetermined training parameter data to obtain a detail classification model; d. Conduct data testing on the obtained detail classification model to quickly distinguish the differences in various image details including image color, image position, image size, image content and image action, and classify the differences in details, determine the optimization method, and obtain the detail classification equation; S4: Input new data and make predictions.