Image classification method and device based on deep learning model, and electronic equipment
By adopting the combination method of large-core row-term convolution and attention module in the image classification model, the limitations of traditional convolution kernels when processing large-scale or high-resolution images are solved, and more efficient and accurate feature extraction and image classification are achieved.
Patent Information
- Application Number
- CN202510153118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional convolution kernel sizes have limitations when processing large-scale or high-resolution images. Smaller convolution kernels cannot fully utilize the global information of the image, while larger convolution kernels will lead to a significant increase in computational complexity, limiting the performance and scalability of deep learning models.
The image classification method of large kernel row-sequence convolution combined with attention module is adopted. Continuous convolution operations are performed by using convolution kernels of size a*a, b*b, 1*c, c*1, 1*d and d*1, and the features are enhanced by the attention module in each step to generate the final feature map for image classification.
This method can better capture the global structure and local features of the image, improve the efficiency and accuracy of feature extraction, adapt to the characteristics of hardware devices, and improve the performance of deep learning models in image classification tasks.
Smart Images

Figure CN120147692A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to an image classification method, apparatus, and electronic device based on a deep learning model. Background Art
[0002] In recent years, deep learning technology, especially deep convolutional neural networks (CNNs), has made remarkable progress in the fields of image processing and computer vision. These networks can achieve human-level performance in various image classification, object detection, and image segmentation tasks based on deep learning models by learning hierarchical features of images. However, with the in-depth research, researchers have found that the improvement of CNN performance is often limited by the choice of convolutional kernel size. Traditional 3x3, 5x5, and 7x7 convolutional kernels are widely used to construct CNN architectures, and these size choices are based on the results of early experiments and theoretical studies, which are considered to be able to effectively capture local features of images and maintain computational efficiency.
[0003] However, these traditional convolutional kernel sizes face certain limitations when processing large-scale or high-resolution images. Smaller convolutional kernels may not be able to fully utilize the global information in the image, while larger convolutional kernels will lead to a significant increase in computational complexity, which poses higher requirements for hardware resources. In addition, with the development of hardware technology, the performance of traditional convolutional kernels on new hardware devices is not always optimal, which limits the performance and scalability of deep learning models in practical applications. Summary of the Invention
[0004] This application provides an image classification method, apparatus, and electronic device based on a deep learning model to solve the problems in the above background art.
[0005] In a first aspect, this application provides an image classification method based on a deep learning model, including:
[0006] Performing a convolution operation on the input image using an a*a convolutional kernel, and then using an attention module to enhance the features to obtain an initial feature map, where a is a positive integer;
[0007] Performing convolution operations on the initial feature map continuously using b*b and a*a convolutional kernels, and using an attention module to enhance the features to obtain an intermediate feature map, where b is a positive integer and b > a;
[0008] Performing a convolution operation on the intermediate feature map first using a 1*c row convolution, then using a c*1 column convolution, and then using an attention module to enhance the features to obtain an enhanced feature map, where c is a positive integer and c > b;
[0009] Perform a convolution operation on the enhanced feature map using a 1*d row convolution first, then perform a convolution operation using a d*1 column convolution, and then use an attention module to enhance the features to obtain the final feature map, where d is a positive integer and d is greater than c;
[0010] Obtain the classification result of the input image through the final feature map.
[0011] Further, before performing the convolution operation on the input image using an a*a convolution kernel and then using an attention module to enhance the features to obtain the initial feature map, it further includes:
[0012] Preprocess the images in the pre-acquired image dataset to obtain the input image.
[0013] Further, the preprocessing of the images in the pre-acquired image dataset to obtain the input image includes:
[0014] Use random cropping to crop the images in the pre-acquired image dataset into a preset size, and perform horizontal flipping and image normalization processing to obtain the input image.
[0015] Further, the performing the convolution operation on the input image using an a*a convolution kernel and then using an attention module to enhance the features to obtain the initial feature map includes:
[0016] Perform a convolution operation on the input image using an a*a convolution kernel, then use an attention module to enhance the features, and obtain the initial feature map through batch normalization, activation function, and max pooling operations.
[0017] Further, the continuously performing convolution operations on the initial feature map using b*b and a*a convolution kernels and using an attention module to enhance the features to obtain the intermediate feature map includes:
[0018] Perform a convolution operation on the initial feature map using a b*b convolution kernel, then use an attention module to enhance the features, and obtain the first feature map through batch normalization, activation function, and max pooling operations;
[0019] Perform a convolution operation on the first feature map using an a*a convolution kernel, and then use an attention module to enhance the features to obtain the intermediate feature map.
[0020] Further, the obtaining the classification result of the input image through the final feature map includes:
[0021] Use multiple fully connected layers on the final feature map to obtain the classification result of the input image.
[0022] Further, a is 3, b is 5, c is 8, and d is 16.
[0023] In a second aspect, the present application provides an image classification device based on a deep learning model, including:
[0024] An initial feature module, configured to perform a convolution operation on an input image using a convolution kernel of a*a, and then use an attention module to enhance the features to obtain an initial feature map, where a is a positive integer;
[0025] An intermediate feature module, configured to continuously perform convolution operations on the initial feature map using convolution kernels of b*b and a*a, and use an attention module to enhance the features to obtain an intermediate feature map, where b is a positive integer and b is greater than a;
[0026] An enhanced feature module, configured to first perform a convolution operation on the intermediate feature map using a 1*c row convolution, then perform a convolution operation using a c*1 column convolution, and then use an attention module to enhance the features to obtain an enhanced feature map, where c is a positive integer and c is greater than b;
[0027] A final feature module, configured to first perform a convolution operation on the enhanced feature map using a 1*d row convolution, then perform a convolution operation using a d*1 column convolution, and then use an attention module to enhance the features to obtain a final feature map, where d is a positive integer and d is greater than c;
[0028] An image classification module, configured to obtain a classification result of the input image through the final feature map.
[0029] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned image classification method based on a deep learning model is implemented.
[0030] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned image classification method based on a deep learning model is implemented.
[0031] The above technical solutions of the present application have the following advantages:
[0032] The image classification method based on a deep learning model provided in the first aspect of the present application uses large-kernel row-column convolution so that the model can capture broader image context information. The large-kernel row-column convolution can cover a larger area in the image, which helps the model understand the global structure of the image. By introducing the attention module into the large-kernel convolution, the model performance is further improved. The attention mechanism enables the model to focus on the key regions in the image, improving the efficiency and accuracy of feature extraction. Through this method, it is possible to better adapt to the characteristics of hardware devices and at the same time improve the performance of the deep learning model in general image classification tasks.
[0033] It can be understood that for the beneficial effects of the above-mentioned second aspect, third aspect and fourth aspect, reference can be made to the relevant descriptions in the above-mentioned first aspect, and details will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 It is a schematic flowchart of the image classification method based on a deep learning model provided by the present application;
[0036] Figure 2 It is a schematic structural diagram of the image classification device based on a deep learning model provided by the present application;
[0037] Figure 3 It is a schematic structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0039] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0040] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0041] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized. "A plurality" means "two or more".
[0042] The objective of this application is to explore a new convolutional network architecture that can better adapt to the characteristics of hardware devices while improving the performance of the model in general image classification tasks. To this end, this application proposes an image classification method, device, and electronic device based on a deep learning model.
[0043] The following will further describe in detail the specific implementation manners of this application in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate this application, but are not used to limit the scope of this application.
[0044] As Figure 1 shown, an image classification method based on a deep learning model provided by an embodiment of this application specifically includes the following steps: performing a convolution operation on an input image using an a*a convolution kernel, and then using an attention module to enhance the features to obtain an initial feature map, where a is a positive integer; performing convolution operations on the initial feature map continuously using b*b and a*a convolution kernels, and using an attention module to enhance the features to obtain an intermediate feature map, where b is a positive integer and b > a; performing a convolution operation on the intermediate feature map first using a 1*c row convolution, then using a c*1 column convolution, and then using an attention module to enhance the features to obtain an enhanced feature map, where c is a positive integer and c > b; performing a convolution operation on the enhanced feature map first using a 1*d row convolution, then using a d*1 column convolution, and then using an attention module to enhance the features to obtain a final feature map, where d is a positive integer and d > c; obtaining the classification result of the input image through the final feature map.
[0045] In some embodiments, before performing the convolution operation on the input image using an a*a convolution kernel and then using an attention module to enhance the features to obtain an initial feature map, it further includes: preprocessing the images in a pre-acquired image dataset to obtain the input image.
[0046] In some embodiments, preprocessing the images in the pre-acquired image dataset to obtain input images includes: using random cropping to crop the images in the pre-acquired image dataset into a preset size, and performing horizontal flipping and image normalization processing to obtain input images.
[0047] In some embodiments, performing a convolution operation on the input image using a convolution kernel of a*a, and then using an attention module to enhance features to obtain an initial feature map includes: performing a convolution operation on the input image using a convolution kernel of a*a, and then using an attention module to enhance features, and obtaining an initial feature map through batch normalization, an activation function, and a max pooling operation.
[0048] In some embodiments, performing convolution operations on the initial feature map consecutively using convolution kernels of b*b and a*a, and using an attention module to enhance features to obtain an intermediate feature map includes: performing a convolution operation on the initial feature map using a convolution kernel of b*b, and then using an attention module to enhance features, and obtaining a first feature map through batch normalization, an activation function, and a max pooling operation; performing a convolution operation on the first feature map using a convolution kernel of a*a, and then using an attention module to enhance features to obtain an intermediate feature map.
[0049] In some embodiments, obtaining the classification result of the input image through the final feature map includes: using multiple fully connected layers on the final feature map to obtain the classification result of the input image.
[0050] In some embodiments, a is 3, b is 5, c is 8, and d is 16.
[0051] Using an image preprocessing module on the CIFAR-10 dataset to obtain the input image of the network. Specifically, using random cropping to crop the input image into a size of 32*32, and further processing it using horizontal flipping and image normalization. Inputting the processed image into a convolution kernel of size 3*3, and then using an attention module (i.e., SE module) to enhance features. Finally, using batch normalization (Batch Normalization), an activation function (ReLU), and a max pooling (Max Pooling) operation to obtain an initial feature map.
[0052] Performing a convolution operation on the initial feature map using a convolution kernel of size 5*5, and then using the SE module to enhance features. Finally, using batch normalization (Batch Normalization), an activation function (ReLU), and a max pooling (MaxPooling) operation to obtain a first feature map, and performing a convolution operation on the first feature map using a convolution kernel of size 3*3, and then using the SE module to enhance features to obtain an intermediate feature map.
[0053] First, a 1×8 row convolution is used to perform a convolution operation on the intermediate feature map, followed by an 8×1 column convolution. Then, the SE module is used to enhance the features to obtain an enhanced feature map. On the enhanced feature map, a 1×16 row convolution is first used for convolution, followed by a 16×1 column convolution. Then, the SE module is used to enhance the features to obtain the final feature map. Three fully connected layers are continuously used on the final feature map to further obtain the final classification.
[0054] The attention module SE is as follows: SENet mainly learns the correlation between channels, filters out the attention for channels, slightly increases the amount of calculation, but has a better effect. Specifically, by processing the feature map obtained by convolution, a one-dimensional vector with the same number of channels as the feature map is obtained as the evaluation score for each channel, and then this score is applied to the corresponding channel respectively to obtain the result.
[0055] The following further elaborates on the present application in combination with specific embodiments.
[0056] Embodiment
[0057] In this embodiment, the effectiveness of the proposed method is evaluated on the CIFAR-10 dataset. CIFAR-10 is a benchmark dataset widely used in the field of computer vision research. This dataset contains 60,000 color images with a size of 32×32 pixels, which are divided into 10 categories, with 6,000 images in each category. These 10 categories are: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, and truck. The dataset is divided into 50,000 training images and 10,000 test images, where both the training set and the test set are evenly distributed, that is, each category has 5,000 images in the training set and 1,000 images in the test set.
[0058] In network training, the stochastic gradient descent optimizer and batch normalization are used as regularization methods. The batch size is 256, the number of epochs is 150, the learning rate is 0.1, the weight decay is 0.0005, and the momentum is 0.9. This embodiment is constructed based on the AlexNet as the benchmark model. Therefore, this embodiment mainly compares the proposed method and the AlexNet network on CIFAR-10. As shown in Tables 1 and 2, combining large-kernel row convolution and the attention mechanism can help the model obtain better performance. In addition, this large-kernel row-column convolution is more suitable for hardware devices.
[0059] Table 1 Comparison of the network structures of AlexNet and the method proposed in this embodiment
[0060]
[0061]
[0062] Table 2 Performance Comparison
[0063]
[0064] The image classification method based on a deep learning model provided by the embodiments of the present application first explores the use of large-kernel row-column convolution so that the model can capture more extensive image context information. The large-kernel row-column convolution can cover a larger area in the image, which helps the model understand the global structure of the image. However, it is not enough to only use large-kernel row-column convolution to capture the global structure, and local information also needs to be captured. For this reason, the present application introduces the attention module SE into the large-kernel convolution to further improve the model performance. The attention mechanism enables the model to focus on the key regions in the image, improving the efficiency and accuracy of feature extraction. Through this method, it can better adapt to the characteristics of hardware devices and at the same time improve the performance of the deep learning model in general image classification tasks.
[0065] Corresponding to the image classification method based on a deep learning model described in the above embodiments, as Figure 2 shown, the embodiments of the present application also provide an image classification device based on a deep learning model. The image classification device based on a deep learning model includes:
[0066] An initial feature module, configured to perform a convolution operation on the input image using a convolution kernel of a*a, and then use an attention module to enhance the features to obtain an initial feature map, where a is a positive integer;
[0067] An intermediate feature module, configured to continuously perform convolution operations on the initial feature map using convolution kernels of b*b and a*a, and use an attention module to enhance the features to obtain an intermediate feature map, where b is a positive integer and b is greater than a;
[0068] An enhanced feature module, configured to first perform a convolution operation on the intermediate feature map using a 1*c row convolution, then perform a convolution operation using a c*1 column convolution, and then use an attention module to enhance the features to obtain an enhanced feature map, where c is a positive integer and c is greater than b;
[0069] A final feature module, configured to first perform a convolution operation on the enhanced feature map using a 1*d row convolution, then perform a convolution operation using a d*1 column convolution, and then use an attention module to enhance the features to obtain a final feature map, where d is a positive integer and d is greater than c;
[0070] An image classification module, configured to obtain a classification result of the input image through the final feature map.
[0071] It should be noted that for the information interaction, execution process, etc. between the above-mentioned modules / units, since they are based on the same concept as the method embodiments of this application, for their specific functions and the technical effects brought about, reference can be specifically made to the method embodiment section, and details will not be elaborated here.
[0072] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above-mentioned system can refer to the corresponding processes in the foregoing method embodiments, and details will not be elaborated here.
[0073] The embodiment of this application also provides an electronic device, such as Figure 3 shown, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the image classification method provided in the first aspect based on the deep learning model.
[0074] In applications, the electronic device may include, but is not limited to, a processor and a memory, Figure 3 which are only examples of the electronic device and do not constitute a limitation on the electronic device. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, input / output devices, network access devices, etc. The input / output device may include a camera, an audio acquisition / playback device, a display screen, etc. The network access device may include a network module for performing wireless communication with external devices.
[0075] In an application, the processor can be a Central Processing Unit (CPU), and the processor can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0076] In an application, in some embodiments, the memory can be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. In other embodiments, the memory can also be an external storage device of the electronic device, for example, a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory can also include both the internal storage unit and the external storage device of the electronic device. The memory is used to store an operating system, application programs, a BootLoader, data, and other programs, such as program codes of computer programs. The memory can also be used to temporarily store data that has been output or will be output.
[0077] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented.
[0078] All or part of the processes in the method of the above embodiments of the present application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to an electronic device, a recording medium, a computer memory, a Read-Only Memory (ROM), a Random Access Memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.
[0079] Those of ordinary skill in the art can realize that the devices and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0080] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. Additionally, the couplings or direct couplings or communication connections shown or discussed among each other can be through some interfaces. The devices are indirectly coupled or communication-connected, and can be in electrical, mechanical, or other forms.
[0081] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. An image classification method based on a deep learning model, characterized in that: include: The input image is convolved with a convolution kernel of a*a, and then the attention module is used to enhance the features to obtain the initial feature map, where a is a positive integer; The initial feature map is convolved using convolution kernels of b*b and a*a successively, and an attention module is used to enhance the feature to obtain an intermediate feature map, where b is a positive integer and b is greater than a; The intermediate feature map is first convolved using a 1*c row convolution, then convolved using a c*1 column convolution, and then the feature is enhanced using an attention module to obtain an enhanced feature map, where c is a positive integer and c is greater than b; The enhanced feature map is first convolved using a 1*d row convolution, then convolved using a d*1 column convolution, and then the feature is enhanced using an attention module to obtain a final feature map, where d is a positive integer and d is greater than c; The classification result of the input image is obtained through the final feature map.
2. The image classification method based on the deep learning model according to claim 1, characterized in that: Before performing a convolution operation on the input image using a convolution kernel of a*a and then using an attention module to enhance features to obtain an initial feature map, the method further includes: The images in the pre-acquired image dataset are preprocessed to obtain input images.
3. The image classification method based on the deep learning model according to claim 2, characterized in that: The preprocessing of the images in the pre-acquired image data set to obtain the input image includes: The images in the pre-acquired image dataset are cropped into a preset size using random cropping, and are horizontally flipped and normalized to obtain an input image.
4. The image classification method based on the deep learning model according to claim 1, characterized in that: The input image is convolved using a convolution kernel of a*a, and then the attention module is used to enhance the features to obtain the initial feature map, including: The input image is convolved with a convolution kernel of a*a, and then the features are enhanced using the attention module. The initial feature map is obtained through batch normalization, activation function and maximum pooling operations.
5. The image classification method based on deep learning model according to claim 1, characterized in that: The method of continuously using convolution kernels of b*b and a*a to perform convolution operations on the initial feature map and using an attention module to enhance features to obtain an intermediate feature map includes: Performing a convolution operation on the initial feature map using a b*b convolution kernel, and then using an attention module to enhance the features, and obtaining a first feature map through batch normalization, activation function, and maximum pooling operations; A convolution operation is performed on the first feature map using a convolution kernel of a*a, and then the attention module is used to enhance the features to obtain an intermediate feature map.
6. The image classification method based on deep learning model according to claim 1, characterized in that: The obtaining the classification result of the input image through the final feature map includes: A plurality of fully connected layers are used on the final feature map to obtain a classification result of the input image.
7. The image classification method based on deep learning model according to claim 1, characterized in that: a is 3, b is 5, c is 8, and d is 16.
8. An image classification device based on a deep learning model, characterized in that: include: The initial feature module is used to perform a convolution operation on the input image using a convolution kernel of a*a, and then use the attention module to enhance the features to obtain the initial feature map, where a is a positive integer; An intermediate feature module, used to continuously perform convolution operations on the initial feature map using convolution kernels of b*b and a*a, and use an attention module to enhance features to obtain an intermediate feature map, where b is a positive integer and b is greater than a; An enhanced feature module is used to first perform a convolution operation on the intermediate feature map using a 1*c row convolution, then perform a convolution operation using a c*1 column convolution, and then use an attention module to enhance the feature to obtain an enhanced feature map, where c is a positive integer and c is greater than b; A final feature module is used to first perform a convolution operation on the enhanced feature map using a 1*d row convolution, then perform a convolution operation using a d*1 column convolution, and then use an attention module to enhance the feature to obtain a final feature map, where d is a positive integer and d is greater than c; An image classification module is used to obtain a classification result of the input image through the final feature map.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the image classification method based on the deep learning model as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image classification method based on the deep learning model as described in any one of claims 1 to 7 is implemented.