Picture classification method and device, electronic equipment and readable storage medium
The image features are processed through the convolution module, cross-stage module and large-core attention mechanism in the preset detection model, and the problem of imbalance in the detection speed and accuracy of target objects in the picture is solved, and efficient and accurate picture classification is achieved.
Patent Information
- Application Number
- CN202510222039.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the detection speed and accuracy of the target object in the picture are unbalanced, the sensor-based method is cost-effective and has low cost-effectiveness, while the deep learning vision algorithm is insufficient real-time.
Using a preset detection model, the image features are extracted through the first convolution module, combined with the preset intersection stage module, the second efficient multi-scale attention mechanism and the preset large-core attention mechanism, the wavelet convolution is used to process the feature vectors to improve the accuracy and efficiency of feature extraction.
It improves the real-time and accuracy of image classification, can quickly and accurately determine the classification probability of target objects in the image, and solves the problem of imbalance in speed and accuracy.
Smart Images

Figure CN120339673A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of image processing, and in particular, to an image classification method, apparatus, electronic device, and readable storage medium. Background Art
[0002] In the prior art, the methods for detecting target objects in images mainly include sensor-based methods and deep learning vision algorithm-based methods. Among them, the sensor-based method has high real-time performance and accuracy in detecting target objects, but the cost is too high and the cost performance is not high. The deep learning vision algorithm-based method has advantages such as high cost performance, but the real-time performance is insufficient, and there is a problem of imbalance between the speed and accuracy of detecting target objects in images. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide an image classification method, apparatus, electronic device, and readable storage medium to solve the problem of imbalance between the speed and accuracy of detecting target objects in images in the prior art.
[0004] In the first aspect of the embodiments of the present disclosure, an image classification method is provided, including:
[0005] Receiving an image to be classified, and extracting first image features of the image to be classified through a first convolutional module of a preset detection model;
[0006] Processing the first image features through a preset cross-stage module in the preset detection model to obtain a first feature vector, where the preset cross-stage module includes a preset residual component, and a first efficient multi-scale attention mechanism is set in the preset residual component;
[0007] Processing the first feature vector through a second efficient multi-scale attention mechanism to obtain a second feature vector;
[0008] Inputting the second feature vector into a preset large kernel attention mechanism to obtain a third feature vector; the preset large kernel attention mechanism is obtained based on wavelet convolution;
[0009] Based on the third feature vector, obtaining the classification probability of the target object in the image to be classified.
[0010] In the second aspect of the embodiments of the present disclosure, an image classification apparatus is provided, including:
[0011] An extraction module, configured to receive an image to be classified, and extract first image features of the image to be classified through a first convolutional module of a preset detection model;
[0012] The first processing module is configured to process the first image feature through a preset cross-stage module in a preset detection model to obtain a first feature vector. The preset cross-stage module includes a preset residual component, and a first efficient multi-scale attention mechanism is set in the preset residual component;
[0013] The second processing module is configured to process the first feature vector through a second efficient multi-scale attention mechanism to obtain a second feature vector;
[0014] The third processing module is configured to input the second feature vector into a preset large-kernel attention mechanism to obtain a third feature vector; the preset large-kernel attention mechanism is obtained based on wavelet convolution;
[0015] The determination module is configured to obtain the classification probability of the target object in the image to be classified based on the third feature vector.
[0016] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0017] In a fourth aspect of the embodiments of the present disclosure, a readable storage medium is provided. The readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0018] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: receiving an image to be classified, extracting the first image feature of the image to be classified through the first convolution module of the preset detection model, processing the first image feature through the preset cross-stage module in the preset detection model to obtain a first feature vector, where the preset cross-stage module includes a preset residual component, and a first efficient multi-scale attention mechanism is set in the preset residual component, processing the first feature vector through a second efficient multi-scale attention mechanism to obtain a second feature vector, inputting the second feature vector into a preset large-kernel attention mechanism to obtain a third feature vector, the preset large-kernel attention mechanism is obtained based on wavelet convolution, and obtaining the classification probability of the target object in the image to be classified according to the third feature vector, improving the speed and real-time performance of determining the classification probability of the target object in the image to be classified, and improving the calculation accuracy, so as to be able to quickly classify the image to be classified and ensure real-time performance and accuracy. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic flowchart of a picture classification method provided by an embodiment of the present disclosure;
[0021] Figure 2 It is a schematic flowchart of a preset processing module in a preset detection model provided by an embodiment of the present disclosure;
[0022] Figure 3 It is a schematic flowchart of a preset residual component provided by an embodiment of the present disclosure;
[0023] Figure 4 It is a schematic flowchart of a preset cross-stage module provided by an embodiment of the present disclosure;
[0024] Figure 5 It is a schematic flowchart of a preset large kernel attention mechanism provided by an embodiment of the present disclosure;
[0025] Figure 6 It is a schematic flowchart of a preset large kernel attention network provided by an embodiment of the present disclosure;
[0026] Figure 7 It is a schematic flowchart of a preset module including a preset large kernel attention mechanism provided by an embodiment of the present disclosure;
[0027] Figure 8 It is a schematic diagram of a picture classification device provided by an embodiment of the present disclosure;
[0028] Figure 9 It is a schematic diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0029] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0030] A picture classification method and device according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0031] Figure 1 It is a schematic flowchart of a picture classification method provided by an embodiment of the present disclosure. Figure 1 The picture classification method can be executed by a terminal device or a server. As Figure 1 shown, the picture classification method includes:
[0032] S101. Receive the image to be classified, and extract the first image feature of the image to be classified through the first convolutional module of the preset detection model.
[0033] Specifically, receive the image to be classified. The image to be classified can be obtained by cropping an initial image captured by a mobile terminal into multiple images according to a preset cropping size. The mobile terminal can include a drone or a camera or other mobile terminals that can be used for taking pictures, and the initial image can be cropped into multiple images to be classified according to the preset cropping size, so that the preset detection model can process the image to be classified more efficiently.
[0034] In addition, input the image to be classified into the preset detection model, and enable the first convolutional module in the preset detection model to process the image to be classified to obtain the first image feature of the image to be classified. Among them, the first convolutional module can include a convolutional layer, a normalization layer, and an activation function.
[0035] Among them, the parameters such as the size, number, and stride of the convolutional kernel corresponding to the convolutional layer in the first convolutional module are all preset in advance, and can perform a convolutional operation on the pixel information in the image to be classified, so as to extract the first image feature of the image. The first image feature contains basic visual information such as the edges and textures of the objects in the image.
[0036] S102. Process the first image feature through the preset cross-stage module in the preset detection model to obtain the first feature vector.
[0037] The preset cross-stage module includes a preset residual component, and a first efficient multi-scale attention mechanism is set in the preset residual component.
[0038] Specifically, process the extracted first image feature through the preset cross-stage module in the preset detection model. The preset cross-stage module is a network structure that can perform complex transformations and fusions on the input features. The preset cross-stage module includes a preset residual component, and the preset residual component is used to solve the problem of gradient disappearance or gradient explosion during the training process of the deep network, so that the preset cross-stage module can learn features more effectively.
[0039] Furthermore, in the preset residual component, a first efficient multi-scale attention mechanism is set. The first efficient multi-scale attention mechanism can simultaneously focus on the information of the first image feature of the image to be classified at different scales. For example, different weights can be given to the local details and the overall structure respectively, so as to extract features more comprehensively and obtain the first feature vector.
[0040] Among them, the preset cross-stage processing module can be CSPX, where X represents the number of iterations and takes a positive integer. That is, the number represented by X indicates the number of preset residual components in the preset residual component. The preset residual component can be a ResUnit component, and the first efficient multi-scale attention mechanism can be EMA.
[0041] S103, process the first feature vector through the second efficient multi-scale attention mechanism to obtain a second feature vector.
[0042] Specifically, process the first feature vector through the second efficient multi-scale attention mechanism. Among them, the second efficient multi-scale attention mechanism can perform weighted processing on features at different scales to further strengthen the part of the first feature vector that is more valuable for the classification task, while suppressing some unimportant information, to obtain a second feature vector.
[0043] Among them, the second efficient multi-scale attention mechanism can be the second EMA. The process of the second EMA processing the first feature vector is to first group the first feature vector, then use parallel sub-networks to extract features of different channels and scales, and finally enrich feature aggregation through cross-space learning. Specifically:
[0044] (1) Feature grouping: For any given input feature map (i.e., the first feature vector in the present disclosure), EMA divides X into G sub-feature groups, and each group learns different semantics. The feature grouping method enables the model to be distributed and processed on more GPU resources, can strengthen the feature learning of semantic regions, and also compresses noise.
[0045] (2) Parallel sub-networks: EMA uses three parallel paths to extract the attention weight descriptors of the grouped feature maps. Two paths are 1x1 branches, and the third path is a 3x3 branch. One-dimensional global average pooling operations are used in the 1x1 branches to encode channel information in two spatial directions respectively, and the 3x3 branch captures multi-scale feature representations through 3x3 convolutions, enabling EMA to not only encode cross-channel information to adjust the importance of different channels, but also retain accurate spatial structure information in the channels.
[0046] (3) Cross - spatial learning: EMA enriches feature aggregation by providing a cross - spatial information aggregation method in different spatial dimension directions. Specifically, the output of the 1x1 branch encodes global spatial information through two - dimensional global average pooling, while the output of the 3x3 branch is directly transformed into the corresponding dimensional shape. Then, these outputs are aggregated through a matrix dot - product operation to generate the first spatial attention map. Finally, the output feature maps within each group are aggregated through the Sigmoid function of the two generated spatial attention weight values, capturing pixel - level pairing relationships and highlighting the global context of all pixels. The final output of EMA has the same size as the input and can be effectively stacked into modern architectures.
[0047] The first feature vector is processed by the parallel sub - structures in the second efficient multi - scale attention mechanism, avoiding more sequential processing and greater depth during the processing, improving the computational efficiency. At the same time, by combining the convolutional branches with different strides in the second EMA, it can effectively capture multi - scale spatial structure information. Through the cross - spatial learning method, the second EMA can capture global context information and improve the discriminative ability of feature representation.
[0048] S104, input the second feature vector into a preset large - kernel attention mechanism to obtain a third feature vector.
[0049] Among them, the preset large - kernel attention mechanism is based on wavelet convolution.
[0050] Specifically, input the second feature vector into the preset large - kernel attention mechanism. Among them, the preset large - kernel attention mechanism is a special attention mechanism based on wavelet convolution. Wavelet convolution can perform multi - resolution analysis on features and capture information of features at different frequencies.
[0051] The preset large - kernel attention mechanism utilizes this characteristic of wavelet convolution to better focus on the important parts in the features, thereby processing the second feature vector to obtain a third feature vector, so that the third feature vector contains feature information crucial for the classification task.
[0052] Among them, wavelet convolution can be denoted as WTConv, and the preset large - kernel attention mechanism can be denoted as WTLKAttention. The process of obtaining the preset large - kernel attention mechanism through wavelet convolution is as follows: filter and down - sample the input low - frequency and high - frequency content through wavelet transform (WT), then perform depth convolution on different frequency maps, and finally use inverse wavelet transform (IWT) to construct the output.
[0053] Each level of the wavelet transform increases the receptive field size of the layer and can be achieved by slightly increasing the number of trainable parameters. That is, each level of the WT cascades frequency decomposition, plus the fixed-size convolutional kernels at each level, makes the number of parameters grow linearly with the number of levels, while the receptive field grows exponentially. The way of obtaining the preset large kernel attention mechanism based on the wavelet convolution of the wavelet transform can better capture low-frequency information, has strong robustness and the ability to strongly respond to the shape of the object.
[0054] S105. Based on the third eigenvector, obtain the classification probability of the target object in the image to be classified.
[0055] Specifically, based on the obtained third eigenvector, through some subsequent processing steps, such as a classifier, etc., the classification probability of the target object in the image to be classified can be determined. For example, if the image to be classified contains multiple objects, by analyzing the third eigenvector, the area ratio or pixel ratio of the target object in the image can be determined, etc., so as to obtain the classification probability information of the target object, providing an important reference basis for the image classification task.
[0056] According to the technical solution provided by the embodiments of the present disclosure, the eigenvector corresponding to the image to be classified can be extracted comprehensively, accurately and efficiently, effectively solving the gradient problem in the training of deep networks. At the same time, the introduction of the efficient multi-scale attention mechanism and the large kernel attention mechanism enables the preset detection model to better focus on the important features at different scales and frequencies, improving the accuracy and robustness of classification.
[0057] In some embodiments, each preset processing module includes a preset cross-stage module, a second efficient multi-scale attention mechanism, and a preset large kernel attention mechanism;
[0058] Based on the third eigenvector, obtaining the classification probability of the target object in the image to be classified includes:
[0059] According to the preset number of iterations and preset iteration parameters, determine the third eigenvector as the input of the next preset processing module, so as to process the input eigenvector through the preset cross-stage module, process the eigenvector output by the preset cross-stage module through the second efficient multi-scale attention mechanism, and process the eigenvector output by the second efficient multi-scale attention mechanism through the preset large kernel attention mechanism to obtain the target eigenvector corresponding to the image to be classified;
[0060] Based on the target eigenvector, determine the classification probability of the target object in the image to be classified.
[0061] Specifically, the preset detection model may include multiple preset processing modules. Each preset processing module includes a preset cross-stage module, a second efficient multi-scale attention mechanism, and a preset large kernel attention mechanism, and the multiple preset processing modules are connected to each other.
[0062] According to the preset number of iterations and preset iteration parameters, the third feature vector is used as the input of the first preset processing module. Each iteration corresponds to a preset processing module, and the preset number of iterations and preset iteration parameters determine the processing method and degree of the feature vector in the module.
[0063] During the process of the preset cross-stage module processing the input feature vector, through feature fusion and transformation operations, the expression ability of the features is further enhanced, and the problem of information loss in the deep network can be effectively solved.
[0064] In addition, the feature vector output by the preset cross-stage module is input into the second efficient multi-scale attention mechanism to simultaneously focus on the information of the feature vector at different scales. By determining the weights at different scales, the feature vector is weighted processed, thereby effectively improving the comprehensiveness and accuracy of the features, and ensuring that the features contain rich information from local details to overall structure.
[0065] Furthermore, the feature vector processed by the second efficient multi-scale attention mechanism is input into the preset large kernel attention mechanism. Since the preset large kernel attention mechanism is based on wavelet convolution, it can perform multi-resolution analysis on the feature vector, capture the information at different frequencies, and further strengthen the part of the feature that is more valuable for the classification task through large kernel convolution operations, while suppressing unimportant information.
[0066] In addition, according to the preset number of iterations and preset iteration parameters, the feature vector is processed by the preset cross-stage module, the second efficient multi-scale attention mechanism, and the preset large kernel attention mechanism to obtain the target feature vector corresponding to the picture to be classified.
[0067] Among them, the preset number of iterations can be 3, and the preset iteration parameters can be used to represent the number of preset residual components included in the preset cross-stage module in each iteration process.
[0068] In addition, according to the target feature vector, the classification probability of the target object in the picture to be classified can be determined through a classifier or other analysis tools. For example, if the picture to be classified contains multiple objects, by analyzing the target feature vector, the area ratio or pixel ratio of the target object in the picture to be classified can be determined, so as to determine the classification probability of the target object.
[0069] Figure 2 It is a schematic flow diagram of a preset processing module in a preset detection model provided by an embodiment of the present disclosure, as Figure 2As shown, the preset detection model includes 4 preset processing modules. That is, after obtaining the third feature vector through the first preset processing module, three more preset processing modules are required to process the third feature vector to obtain the target feature vector. That is, the preset number of iterations is 3 times, and the preset iteration parameters are 2, 2, 8, and 4 in sequence, indicating that the first preset processing module includes two preset residual components, and so on.
[0070] Among them, CSPX is used to represent the preset cross-stage module, X is used to represent the preset iteration parameter corresponding to this preset processing module, EMA2 is used to represent the second efficient multi-scale attention mechanism, and WTLKACSP is used to represent the processing module including the preset large kernel attention mechanism.
[0071] According to the technical solution provided by the embodiments of the present disclosure, the process of feature extraction and processing can be gradually optimized. Through the synergistic effect of the preset cross-stage module, the second efficient multi-scale attention mechanism, and the preset large kernel attention mechanism, the features are effectively enhanced and optimized at different stages. The finally generated target feature vector can more accurately reflect the information of the target object in the image to be classified, thereby improving the accuracy and reliability of the target object proportion calculation, improving the quality of the features, enhancing the adaptability of the preset detection model to complex images, and making the entire image classification process more efficient and robust.
[0072] In some embodiments, the preset cross-stage module further includes a second convolutional module and a third convolutional module;
[0073] Processing the first image feature through the preset cross-stage module in the preset detection model to obtain the first feature vector, including:
[0074] Processing the first image feature through the second convolutional module to obtain the first initial feature;
[0075] Processing the first initial feature through the third convolutional module to obtain the second initial feature;
[0076] Processing the second initial feature through the preset residual component to obtain the third initial feature;
[0077] Based on the first initial feature and the third initial feature, obtain the first feature vector.
[0078] Specifically, in addition to including the preset residual component, the preset cross-stage module further includes a second convolutional module and a third convolutional module,
[0079] Input the first image feature extracted from the first convolutional module into the second convolutional module. The second convolutional module consists of multiple convolutional layers, and the parameters of the convolutional layers, such as the kernel size, number, stride, etc., are preset. The second convolutional module is used to further extract and refine the features of the first image feature, capture deeper feature information. Through convolutional operations, the second convolutional module can extract information such as edges, textures, and shapes in the features and transform them into more advanced feature representations. Determine the first image feature processed by the second convolutional layer as the first initial feature.
[0080] In addition, input the first initial feature into the third convolutional module. The third convolutional module also consists of multiple convolutional layers, and its parameters are also preset. The third convolutional module is used to further refine and enhance the feature representation ability on the basis of the second convolutional module, and can perform deeper feature extraction on the first initial feature to further capture the detailed information in the features. Through the processing of the third convolutional module, a more advanced feature representation, that is, the second initial feature, can be obtained.
[0081] Furthermore, input the second initial feature into the preset residual component. The design inspiration of the preset residual component comes from the Residual Network (ResNet), and the problem of gradient disappearance or gradient explosion in the training of deep networks can be solved by introducing residual connections. In the preset residual component, the input second initial feature will go through a series of convolutional layers for feature extraction, and at the same time, the original information of the input feature will be retained. Through the connection of the preset residual component, a more stable feature representation, that is, the third initial feature, is obtained.
[0082] In addition, generate the first feature vector through the first initial feature and the third initial feature. The first initial feature and the third initial feature can be weighted summed or concatenated through feature fusion to obtain the first feature vector, which can make full use of the feature information extracted by the second convolutional module and the preset residual component, and at the same time retain the important details in the first initial feature, so that the obtained first feature vector contains rich feature information and also has better stability and robustness.
[0083] According to the technical solution provided by the embodiments of the present disclosure, through the above multi-stage feature processing process, the preset cross-stage module can gradually optimize the feature extraction and expression ability. The synergistic effect of the second convolutional module and the third convolutional module enables the features to be effectively refined and enhanced at different stages, while the preset residual component ensures the stability and robustness of the feature extraction process, so that the generated first feature vector not only contains rich feature information, but also has better stability and adaptability, and can significantly improve the performance and accuracy of the preset detection model.
[0084] In some embodiments, the preset residual component further includes two preset convolutional modules;
[0085] Processing the second initial feature through the preset residual component to obtain a third initial feature, including:
[0086] Performing feature extraction on the second initial feature through two preset convolutional modules to obtain a first initial processed feature;
[0087] Processing the first initial processed feature through the first efficient multi-scale attention mechanism to obtain a second initial processed feature;
[0088] Adding the second initial feature and the second initial processed feature to obtain a third initial feature.
[0089] Figure 3 is a schematic flowchart of a preset residual component provided by an embodiment of the present disclosure. As Figure 3 shown, the order of processing the input features in the preset residual component is two preset convolutional modules CBA2 and CBA3, the first efficient multi-scale attention mechanism EMA1, and the addition module add.
[0090] Specifically, the preset residual component further includes two preset convolutional modules, that is, in addition to the first efficient multi-scale attention mechanism in the preset residual component, there are also two preset convolutional modules.
[0091] Performing feature extraction on the second initial feature through two preset convolutional modules. The preset convolutional module is used to further extract and refine the input features, that is, processing the second initial feature through the first preset convolutional module and determining its output as the input of another preset convolutional module, so that another preset convolutional module performs feature extraction to obtain a first initial processed feature.
[0092] In addition, inputting the first initial processed feature into the first efficient multi-scale attention mechanism, where the first efficient multi-scale attention mechanism can be EMA. EMA is suitable for processing complex image features because it can capture both local details and overall structure information at the same time. After being processed by the first efficient multi-scale attention mechanism, a more optimized feature representation, that is, a second initial processed feature, is obtained.
[0093] Furthermore, adding the second initial feature and the second initial processed feature to obtain a third initial feature can ensure that in a deep network, feature information will not be lost, and at the same time, it can also solve the problems of gradient disappearance or gradient explosion, so that the obtained third initial feature has a more stable and rich feature representation.
[0094] According to the technical solution provided by the embodiments of the present disclosure, the preset residual component can effectively process the input features and solve the gradient problem in the training of deep networks through residual connections. The synergistic effect of the two preset convolution modules and the introduction of the first efficient multi-scale attention mechanism enable the features to be effectively refined and enhanced at different stages, making the obtained third initial feature contain rich feature information, and also having better stability and adaptability, which can significantly improve the performance and accuracy of the preset detection model.
[0095] In some embodiments, before obtaining the first feature vector based on the first initial feature and the third initial feature, it further includes:
[0096] Obtaining the number of components of the preset residual component included in the current preset cross-stage module;
[0097] When the number of components is multiple, determining the output of the previous preset residual component as the input of the next preset residual component, so that the second initial feature is processed by multiple preset residual components, and determining the output of the last preset residual component as the third initial feature.
[0098] Specifically, determining the number of components of the preset residual component included in the current preset cross-stage module, that is, the preset iteration parameter mentioned in the above embodiments of the present disclosure. This parameter can determine the subsequent feature processing flow and complexity, and the number of components can be adjusted according to the design requirements of the preset detection model. For example, for a more complex image classification task, more components of the preset residual component can be used to extract richer features. It can clarify the structure of the preset cross-stage module, provide a basis for the subsequent feature processing flow, and ensure the clarity and efficiency of the feature processing logic.
[0099] In addition, taking the output of the first preset residual component as the input of the second preset residual component, and the output of the second preset residual component as the input of the third preset residual component, and so on. The cascaded processing method can ensure that the second initial feature is gradually optimized and enhanced in multiple preset residual components.
[0100] For example, when the current preset cross-stage module includes three preset residual components:
[0101] The first preset residual component: receives the second initial feature as input, and after being processed by two internal preset convolution modules and the first efficient multi-scale attention mechanism, outputs an optimized feature representation.
[0102] The second preset residual component: takes the output of the first preset residual component as input, and through a similar processing process again, further optimizes the feature representation.
[0103] The third preset residual component: taking the output of the second preset residual component as the input, and continuing to perform feature optimization processing.
[0104] In addition, through the cascaded processing of multiple preset residual components, the output of the last preset residual component is determined as the third initial feature. Among them, the third initial feature has undergone multiple optimization processes and contains richer and more stable feature information, providing a high-quality feature basis for the generation of the subsequent first feature vector.
[0105] The cascaded processing can ensure that features are gradually optimized and enhanced in multiple preset residual components, improve the quality and stability of features, and can also make full use of the processing capabilities of each preset residual component to avoid the loss of feature information.
[0106] Figure 4 It is a schematic flowchart of a preset cross-stage module provided by an embodiment of the present disclosure. As Figure 4 shown, the preset cross-stage module includes five preset convolution modules, namely CBA4, CBA5, CBA6, CBA7, and CBA8, and also includes at least one preset residual component. The preset residual components included in the preset cross-stage module can be generally represented as ResUnit. The preset iteration parameter corresponding to the preset cross-stage module represents the number of components with preset residual components in ResUnit. The preset iteration parameter and the number of components can be represented by X, and X is a positive integer.
[0107] Further, the input of the preset cross-stage module is the first image feature. The first image feature is processed by CBA4 to obtain the first image intermediate feature. The first image intermediate feature is input into CBA5, the preset residual component, and CBA6. And the output of CBA5 is determined as the input of the preset residual component, the output of the preset residual component is determined as the input of CBA6, and the output of CBA6 is determined as the second image intermediate feature. At the same time, the first image intermediate feature is input into CBA7 to obtain the third image intermediate feature. The second image intermediate feature and the third image intermediate feature are processed by the connection function Concat, and the processing result is input into CBA8 to obtain the first feature vector.
[0108] According to the technical solution provided by the embodiment of the present disclosure, the preset cross-stage module can perform multiple optimization processes on the input second initial feature. Each preset residual component can refine and enhance the feature through the internal convolution module and attention mechanism. And multiple preset residual components gradually process the second initial feature, enabling the second initial feature to be gradually accumulated and optimized in multiple stages, so that the generated third initial feature includes rich feature information and also has better stability and robustness.
[0109] In some embodiments, the preset large kernel attention mechanism includes depthwise separable wavelet convolution, depthwise dilated separable convolution, and a preset convolutional layer;
[0110] Inputting the second feature vector into the preset large kernel attention mechanism to obtain a third feature vector includes:
[0111] Inputting the second feature vector into the depthwise separable wavelet convolution to obtain a first separated feature;
[0112] Processing the first separated feature through the depthwise dilated separable convolution to obtain a second separated feature;
[0113] Performing feature extraction on the second separated feature through the preset convolutional layer to obtain a third separated feature;
[0114] Multiplying the second feature vector by the third separated feature to obtain the third feature vector.
[0115] Specifically, the preset large kernel attention mechanism is used to further optimize and enhance the expressive ability of the feature vector. It can extract and strengthen the important information in the feature, providing a higher-quality feature representation for subsequent object detection or classification tasks. Among them, the preset large kernel attention mechanism mainly includes three parts: depthwise separable wavelet convolution, depthwise dilated separable convolution kernel, and a preset convolutional layer.
[0116] Inputting the second feature vector into the depthwise separable wavelet convolution in the preset large kernel attention mechanism. Among them, the depthwise separable wavelet convolution is an efficient convolution operation that combines the advantages of wavelet transform and depthwise separable convolution. The wavelet transform can perform multi-resolution analysis on the feature, capturing information at different frequencies, while the depthwise separable convolution reduces the computational amount and the number of parameters by separating the channel and spatial dimensions of the convolution kernel. The depthwise separable wavelet convolution can effectively extract the important information in the feature while maintaining computational efficiency. After processing the second feature vector through the depthwise separable wavelet convolution, a new feature representation, that is, the first separated feature, is obtained.
[0117] In addition, inputting the first separated feature into the depthwise dilated separable convolution, and determining the first separated feature processed by the depthwise dilated separable convolution as the second separated feature. Among them, the depthwise dilated separable convolution is a convolution operation designed to further enhance the expressive ability of the feature by introducing a collision mechanism. Among them, the collision mechanism can capture the complex relationships between features by simulating the interaction between features, thereby extracting richer feature information.
[0118] In addition, the second separation feature is input into a preset convolutional layer for feature extraction to obtain a third separation feature. The preset convolutional layer is composed of multiple convolutional layers, which are used to capture information such as edges, textures, and shapes in the second separation feature through convolutional operations and transform them into a higher-level feature representation, improving the stability and accuracy of the third separation feature.
[0119] In addition, an element-wise multiplication operation is performed on the second feature vector and the third separation feature to obtain a third feature vector. The process of multiplying the second feature vector by the third separation feature can combine the important information in the above two feature vectors with the optimized feature information to obtain a more comprehensive and stable feature representation, that is, the third feature vector. The third feature vector includes rich feature information and also has better stability and adaptability.
[0120] Figure 5 FIG. Figure 5 As shown, the preset large kernel attention mechanism includes a depthwise separable wavelet convolution, a depthwise dilated separable convolution, and a preset convolutional layer. Among them, Input represents the input feature vector, that is, the second feature vector. WTDw-Conv is used to represent the depthwise separable wavelet convolution. WTDw-D-Conv is used to represent the depthwise dilated separable convolution. The wavelet convolution WT-Conv in WTDw-Conv and WTDw-D-Conv is Wavelet Convolutions. Conv is used to represent the preset convolutional layer. Output represents the output feature vector, that is, the third feature vector.
[0121] Among them, the kernel sizes of WT-DwConv and WTDw-D-Conv are different in different stage modules of the deep learning network architecture CSPDarknet. The dilation rate of WTDw-D-Conv is also different in different stages, and there is a case where it is 0.
[0122] Furthermore, in the process of obtaining the third feature vector through the preset large kernel attention mechanism, further processing is required.
[0123] Figure 6 FIG. Figure 7 FIG. Figure 7As shown, after the preset module receives the initial feature vector, it processes the initial feature vector in the order of the first normalization layer and the attention network to obtain the first feature vector, and performs weighted averaging on the initial feature vector and the first feature vector according to the preset weight to obtain the second feature vector. The second feature vector is processed in the order of the second normalization layer and the fully connected layer to obtain the third feature vector, and weighted averaging is performed on the second feature vector and the third feature vector according to the preset weight to obtain the fourth feature vector.
[0124] Among them, the core of the attention network is the preset large kernel attention mechanism as Figure 5 shown, namely WTLKAttention. The processing flow of the attention network is as Figure 7 shown. 1×1Conv represents a convolutional layer with a stride of 1, DiTAC represents a trainable high-expression activation function, α represents the weight corresponding to the initial feature vector, and β represents the weight corresponding to the first feature vector.
[0125] Among them, DiTAC can be represented by the following formula:
[0126]
[0127] Among them, φ(x) is the cumulative distribution function of the standard normal distribution, and T θ (x) is a learnable continuous piecewise affine transformation (CPAB). This transformation can achieve more complex non-linear deformations, can learn more complex function shapes, improve the expression ability of the preset detection model, and thus improve the detection accuracy of the preset detection model.
[0128] Among them, Figure 7 BN1 in represents the first standard layer, BN2 represents the second standard layer, FFN represents the fully connected layer, and 3×3Conv represents a convolutional layer with a stride of 3.
[0129] Furthermore, Figure 7 the preset module including the preset large kernel attention mechanism represented by can be denoted as WTLKABlock, that is Figure 6 the WTLKABlock shown in Figure 6 The preset large kernel attention network shown also includes two preset convolutional modules, namely CBA9 and CBA10. The second feature vector is input into Figure 6 CBA9 in, and through the processing of the preset large kernel attention mechanism, the first intermediate feature vector is obtained. The second feature vector is input into CBA10 to obtain the second intermediate feature vector. The first intermediate feature vector and the second intermediate feature vector are concatenated through the connection function Concat to obtain the third feature vector.
[0130] According to the technical solution provided by the embodiments of the present disclosure, the preset large-kernel attention mechanism can perform multi-stage optimization processing on the input second feature vector. The collaborative effect of the depthwise separable wavelet convolution, the depthwise dilated separable convolution, and the preset convolutional layer in the preset large-kernel attention mechanism enables the second feature vector to be effectively refined and enhanced at different stages, improving the performance and accuracy of the preset detection model.
[0131] In some embodiments, the picture to be classified is a picture of an agricultural planting area, the target object includes the proportion of the water accumulation area in the current agricultural planting area in the area of the current agricultural planting area, and the classification probability includes the classification probability of having water accumulation;
[0132] After obtaining the classification probability of the target object in the picture to be classified based on the third feature vector, it further includes:
[0133] When the classification probability of having water accumulation is greater than or equal to the preset water accumulation threshold, determine the agricultural planting area corresponding to the current picture to be classified as a water-accumulated planting area;
[0134] When the classification probability of having water accumulation is less than the preset water accumulation threshold, determine the agricultural planting area corresponding to the current picture to be classified as a non-water-accumulated planting area.
[0135] Specifically, the picture to be classified can be a picture of an agricultural planting area, the target object is the proportion of the water accumulation area in the current agricultural planting area in the area of the current agricultural planting area, and the classification probability includes the classification probability of having water accumulation. Among them, detecting the water accumulation area in the agricultural planting area is of great significance for agricultural management and disaster warning, and can help farmers take timely measures to reduce losses.
[0136] When the classification probability of having water accumulation is greater than or equal to the preset water accumulation threshold, it indicates that the water accumulation situation in the current agricultural planting area is relatively serious and may cause greater harm to the crops. The agricultural planting area corresponding to the current picture to be classified can be determined as a water-accumulated planting area.
[0137] Furthermore, when the current agricultural planting area is a water-accumulated planting area, corresponding alarms or notifications can be triggered, so as to timely remind farmers or agricultural managers to take emergency measures, such as draining water or transferring crops, to reduce the negative impact of water accumulation on agricultural production.
[0138] When the classification probability of having water accumulation is less than the preset water accumulation threshold, it indicates that the water accumulation situation in the current agricultural planting area is relatively light, that is, the water accumulation area in the current agricultural planting area has not reached the level that requires immediate emergency measures. The agricultural planting area corresponding to the current picture to be classified is determined as a non-water-accumulated planting area.
[0139] In another embodiment, the classification probabilities output by the preset detection model include a water accumulation classification probability and a non-water accumulation classification probability. The larger value of the water accumulation classification probability and the non-water accumulation classification probability is determined as the intermediate classification probability. When the intermediate classification probability is the water accumulation classification probability and the water accumulation classification probability is greater than or equal to the preset water accumulation threshold, the agricultural planting area corresponding to the current picture to be classified is determined as a water accumulation planting area; when the intermediate classification probability is the non-water accumulation probability, the agricultural planting area corresponding to the current picture to be classified is determined as a non-water accumulation planting area.
[0140] Further, when the intermediate classification probability is the water accumulation probability and is less than the preset water accumulation threshold, the real-time monitoring mode can be entered to continuously monitor the classification probability of the target object in the agricultural planting area corresponding to the current picture to be classified. Through real-time monitoring, the change of the water accumulation area can be timely detected, and farmers can be reminded to take corresponding measures when necessary, which helps to make preparations in advance before the water accumulation situation deteriorates and reduce potential losses.
[0141] According to the technical solution provided by the embodiment of the present disclosure, different measures can be taken based on the size relationship between the classified water accumulation probability and the preset water accumulation threshold, so as to timely respond to serious water accumulation situations, and at the same time, over-interference with minor water accumulation situations can be avoided, improving the monitoring efficiency, providing more accurate decision-making support for agricultural managers, and helping to reduce agricultural losses caused by water accumulation.
[0142] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present disclosure, which will not be elaborated one by one here.
[0143] The following is an embodiment of the device of the present disclosure, which can be used to execute the method embodiment of the present disclosure. For details not disclosed in the embodiment of the device of the present disclosure, please refer to the method embodiment of the present disclosure.
[0144] Figure 8 is a schematic diagram of a picture classification device provided by an embodiment of the present disclosure. As Figure 8 shown, the picture classification device includes: an extraction module 801, a first processing module 802, a second processing module 803, a third processing module 804, and a determination module 805, where:
[0145] The extraction module 801 is configured to receive a picture to be classified and extract first picture features of the picture to be classified through a first convolutional module of a preset detection model;
[0146] The first processing module 802 is configured to process the first picture features through a preset cross-stage module in the preset detection model to obtain a first feature vector. The preset cross-stage module includes a preset residual component, and a first efficient multi-scale attention mechanism is set in the preset residual component;
[0147] The second processing module 803 is configured to process the first feature vector through a second efficient multi-scale attention mechanism to obtain a second feature vector;
[0148] The third processing module 804 is configured to input the second feature vector into a preset large kernel attention mechanism to obtain a third feature vector; the preset large kernel attention mechanism is obtained based on wavelet convolution;
[0149] The determination module 805 is configured to obtain the classification probability of the target object in the picture to be classified based on the third feature vector.
[0150] In some embodiments, each preset processing module includes a preset cross-stage module, a second efficient multi-scale attention mechanism, and a preset large kernel attention mechanism; the determination module 805 is configured to:
[0151] According to the preset number of iterations and preset iteration parameters, determine the third feature vector as the input of the next preset processing module, so as to process the input feature vector through the preset cross-stage module, process the feature vector output by the preset cross-stage module through the second efficient multi-scale attention mechanism, and process the feature vector output by the second efficient multi-scale attention mechanism through the preset large kernel attention mechanism to obtain the target feature vector corresponding to the picture to be classified;
[0152] Based on the target feature vector, determine the classification probability of the target object in the picture to be classified.
[0153] In some embodiments, the preset cross-stage module further includes a second convolution module and a third convolution module; the extraction module 801 is configured to:
[0154] Process the first picture feature through the second convolution module to obtain a first initial feature;
[0155] Process the first initial feature through the third convolution module to obtain a second initial feature;
[0156] Process the second initial feature through a preset residual component to obtain a third initial feature;
[0157] Based on the first initial feature and the third initial feature, obtain a first feature vector.
[0158] In some embodiments, the preset residual component further includes two preset convolution modules; the extraction module 801 is configured to:
[0159] Perform feature extraction on the second initial feature through two preset convolution modules to obtain a first initial processed feature;
[0160] Process the first initial processed feature through a first efficient multi-scale attention mechanism to obtain a second initial processed feature;
[0161] Add the second initial feature and the second initial processing feature to obtain a third initial feature.
[0162] In some embodiments, before the extraction module 801 obtains the first feature vector based on the first initial feature and the third initial feature, it is further configured to:
[0163] Obtain the number of components of the preset residual component included in the current preset cross-stage module;
[0164] In the case where the number of components is multiple, determine the output of the previous preset residual component as the input of the next preset residual component, so that the second initial feature is processed by multiple preset residual components, and determine the output of the last preset residual component as the third initial feature.
[0165] In some embodiments, the preset large kernel attention mechanism includes depthwise separable wavelet convolution, depthwise dilated separable convolution, and a preset convolutional layer; the third processing module 804 is configured to:
[0166] Input the second feature vector into the depthwise separable wavelet convolution to obtain a first separated feature;
[0167] Process the first separated feature through the depthwise dilated separable convolution to obtain a second separated feature;
[0168] Extract features from the second separated feature through the preset convolutional layer to obtain a third separated feature;
[0169] Multiply the second feature vector by the third separated feature to obtain a third feature vector.
[0170] In some embodiments, the picture to be classified is a picture of an agricultural planting area, and the target object includes the waterlogging area; after the determination module 805 obtains the classification probability of the target object in the picture to be classified based on the third feature vector, it is further configured to:
[0171] When the waterlogging area is greater than or equal to the first preset threshold, determine the part of the agricultural planting area corresponding to the current picture to be classified as the waterlogged part;
[0172] When the waterlogging area is less than the first preset threshold and greater than or equal to the second preset threshold, real-time monitor the waterlogging area of the agricultural planting area corresponding to the current picture to be classified;
[0173] When the waterlogging area is less than or equal to the second preset threshold, ignore the part of the agricultural planting area corresponding to the current picture to be classified.
[0174] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0175] Figure 9 is a schematic diagram of the electronic device 9 provided by an embodiment of the present disclosure. As Figure 9 shown, the electronic device 9 of this embodiment includes: a processor 901, a memory 902, and a computer program 903 stored in the memory 902 and executable on the processor 901. When the processor 901 executes the computer program 903, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 901 executes the computer program 903, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0176] The electronic device 9 may be a desktop computer, a notebook, a palm computer, a cloud server, and other electronic devices. The electronic device 9 may include, but is not limited to, the processor 901 and the memory 902. Those skilled in the art can understand that Figure 9 merely examples of the electronic device 9, and do not constitute a limitation on the electronic device 9. It may include more or fewer components than those shown in the figure, or different components.
[0177] The processor 901 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0178] The memory 902 may be an internal storage unit of the electronic device 9. For example, the hard disk or memory of the electronic device 9. The memory 902 may also be an external storage device of the electronic device 9. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 9. The memory 902 may also include both the internal storage unit and the external storage device of the electronic device 9. The memory 902 is used to store computer programs and other programs and data required by the electronic device.
[0179] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0180] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium (such as a computer-readable storage medium). Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0181] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.
Claims
1. A method for image classification, characterized in that, Including: Receiving a picture to be classified, and extracting first picture features of the picture to be classified through a first convolutional module of a preset detection model; Processing the first picture features through a preset cross-stage module in the preset detection model to obtain a first feature vector, where the preset cross-stage module includes a preset residual component, and a first efficient multi-scale attention mechanism is set in the preset residual component; Processing the first feature vector through a second efficient multi-scale attention mechanism to obtain a second feature vector; Inputting the second feature vector into a preset large-kernel attention mechanism to obtain a third feature vector; the preset large-kernel attention mechanism is obtained based on wavelet convolution; Based on the third feature vector, obtaining a classification probability of an object in the picture to be classified.
2. The method according to claim 1, characterized in that, Each preset processing module includes a preset cross-stage module, a second efficient multi-scale attention mechanism, and a preset large-kernel attention mechanism; Based on the third feature vector, obtaining a classification probability of an object in the picture to be classified includes: According to a preset number of iterations and preset iteration parameters, determining the third feature vector as an input of a next preset processing module, so as to process the input feature vector through the preset cross-stage module, process the feature vector output by the preset cross-stage module through the second efficient multi-scale attention mechanism, and process the feature vector output by the second efficient multi-scale attention mechanism through the preset large-kernel attention mechanism to obtain a target feature vector corresponding to the picture to be classified; Based on the target feature vector, determining a classification probability of an object in the picture to be classified.
3. The method according to claim 1, wherein The preset cross-stage module further includes a second convolutional module and a third convolutional module; Processing the first picture features through a preset cross-stage module in the preset detection model to obtain a first feature vector, including: Processing the first picture features through the second convolutional module to obtain first initial features; Processing the first initial features through the third convolutional module to obtain second initial features; Processing the second initial features through the preset residual component to obtain third initial features; Based on the first initial features and the third initial features, obtaining a first feature vector.
4. The method according to claim 3, wherein The preset residual component further includes two preset convolutional modules; Processing the second initial features through the preset residual component to obtain third initial features, including: Performing feature extraction on the second initial features through the two preset convolutional modules to obtain first initial processed features; Processing the first initial processed features through the first efficient multi-scale attention mechanism to obtain second initial processed features; Adding the second initial features and the second initial processed features to obtain the third initial features.
5. The method according to claim 3, characterized in that, Before obtaining the first feature vector based on the first initial features and the third initial features, further including: Obtaining the number of components of the preset residual component included in the current preset cross-stage module; In the case where the number of the components is multiple, the output of the previous preset residual component is determined as the input of the next preset residual component, so that the second initial feature is processed by multiple preset residual components, and the output of the last preset residual component is determined as the third initial feature.
6. The method according to claim 1, wherein The preset large kernel attention mechanism includes depthwise separable wavelet convolution, depthwise dilated separable convolution, and a preset convolution layer; Inputting the second feature vector into the preset large kernel attention mechanism to obtain a third feature vector includes: Inputting the second feature vector into the depthwise separable wavelet convolution to obtain a first separated feature; Processing the first separated feature through the depthwise dilated separable convolution to obtain a second separated feature; Performing feature extraction on the second separated feature through the preset convolution layer to obtain a third separated feature; Multiplying the second feature vector by the third separated feature to obtain the third feature vector.
7. The method according to claim 1, characterized in that, The to-be-classified picture is a picture of an agricultural planting area, the target object includes the proportion of the water accumulation area in the current agricultural planting area to the area of the current agricultural planting area, and the classification probability includes a water accumulation classification probability; After obtaining the classification probability of the target object in the to-be-classified picture based on the third feature vector, it further includes: When the water accumulation classification probability is greater than or equal to a preset water accumulation threshold, the agricultural planting area corresponding to the current to-be-classified picture is determined as a water accumulation planting area; When the water accumulation classification probability is less than the preset water accumulation threshold, the agricultural planting area corresponding to the current to-be-classified picture is determined as a non-water accumulation planting area.
8. An image classification device, characterized in that, It includes: An extraction module, configured to receive a to-be-classified picture and extract a first picture feature of the to-be-classified picture through a first convolution module of a preset detection model; A first processing module, configured to process the first picture feature through a preset cross-stage module in the preset detection model to obtain a first feature vector, the preset cross-stage module includes a preset residual component, and a first efficient multi-scale attention mechanism is set in the preset residual component; A second processing module, configured to process the first feature vector through a second efficient multi-scale attention mechanism to obtain a second feature vector; A third processing module, configured to input the second feature vector into a preset large kernel attention mechanism to obtain a third feature vector; the preset large kernel attention mechanism is obtained based on wavelet convolution; A determination module, configured to obtain the classification probability of the target object in the to-be-classified picture based on the third feature vector.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.