Image classification method, system, device and medium with explainability
Patent Information
- Application Number
- CN202611131665.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-08-28
AI Technical Summary
[0003]但是,深度神经网络的内部工作机理是不透明的,神经网络的可解释性至今仍然是一个世界性难题
1、本申请通过根据子块内像素亮度差异区分平地区域和非平地区域,并分别进行颜色聚类和形状聚类,且在聚类时正负样本各自独立进行,使得特征提取过程不再依赖黑盒式的反向传播更新,而是基于明确的物理意义和类别先验知识构建特征基;这种机制使得模型的每一个特征节点都具有清晰的语义解释,解决了深度神经网络内部机理不透明的问题,同时正负样本独立聚类增强了特征对特定类别的判别力,提高了模型的可控性和针对性。
Smart Images

Figure CN122657618A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and image processing technology, and more specifically, to an interpretable image classification method, system, device, and medium. Background Technology
[0002] Deep neural networks (CNNs) are currently the most common method for image classification. Since their emergence in the field, this technology has developed rapidly. Over the past decade, it has made significant progress in tasks such as image classification, visual object detection, pixel segmentation, visual object tracking, image generation, image object matching, and large-scale image-text models. Furthermore, deep neural networks have also shone brightly in various fields, including speech recognition, text and language processing, physical signal analysis, biomedicine, large language models, human-computer interaction, and game AI, becoming a core foundational technology in the field of artificial intelligence today.
[0003] However, the internal workings of deep neural networks are opaque, and the interpretability of neural networks remains a global challenge. This prevents users from accessing the internal workings of the network model file for optimization, and hinders flexible modifications, deletions, and enhancements. This limitation restricts the flexibility of research and development and deployment. Because of this opacity, deep neural network-based artificial intelligence technologies are vulnerable to security breaches. Even slight noise interference can cause drastic changes in recognition results, providing opportunities for malicious actors to attack the network model.
[0004] Meanwhile, this technology exhibits significant randomness during training, making it prone to getting trapped in local optima. It demands considerable skill from developers, often requiring repeated adjustments to hyperparameters and iterative training to achieve satisfactory results. This drawback increases development costs and the difficulty of use. Furthermore, deep neural network models contain numerous redundant parameters; many layers, neurons, and connections are redundant and inefficient. This not only increases the storage size of the network model but also wastes valuable computational resources, leading to increased power consumption, higher costs, and reduced deployment scope and ease of use.
[0005] In this context, exploring and inventing a new image classification method with transparent internal mechanisms becomes very valuable and meaningful. Summary of the Invention
[0006] In view of the deficiencies in the prior art, the purpose of this invention is to provide an image classification method, system, device and medium with interpretability.
[0007] An interpretable image classification method according to the present invention includes: Step S1: Obtain sample sub-blocks of the sample image, and classify the sample sub-blocks into flat areas and non-flat areas based on the difference between the maximum and minimum brightness values of the pixels within the sample sub-blocks; Step S2: Perform color clustering on sample sub-blocks belonging to the flat area and shape clustering on sample sub-blocks belonging to the non-flat area. During clustering, positive and negative sample images are performed independently to obtain color sub-cluster centers and shape sub-cluster centers. Step S3: Obtain the target sub-block of the target image, and classify the target sub-block into flat area and non-flat area based on the difference between the maximum and minimum brightness values of the pixels within the target sub-block; Step S4: Calculate the responsivity of the target sub-block based on the color sub-cluster centers and the shape sub-cluster centers; Step S5: Measure the response into integers and use city distance to calculate the distance between feature vectors to generate the feature vector of the target image; Step S6: Classify the target image based on the feature vector.
[0008] Preferably, the step of dividing the sample sub-block into flat and non-flat areas based on the difference between the maximum and minimum brightness values of pixels within the sample sub-block includes: When the difference between the maximum and minimum brightness values of all pixels in the sample sub-block is less than a preset threshold, the sample sub-block is determined as a flat area. When the difference is greater than or equal to the preset threshold, the sample sub-block is determined to be a non-flat area; The step of performing color clustering on sample sub-blocks belonging to the flat area includes: Calculate the average color value of the sample sub-blocks belonging to the flat area to obtain a color feature vector, and perform color clustering based on the color feature vector; The shape clustering of sample sub-blocks belonging to the non-flat area includes: Brightness normalization is performed on sample sub-blocks belonging to the non-flat area to obtain shape feature vectors, and shape clustering is performed based on the shape feature vectors.
[0009] Preferably, the method further includes a feature extraction and clustering process for multi-layer sub-blocks: In the current layer sub-block, the current layer sub-block contains multiple previous layer sub-blocks; The feature vector corresponding to the previous layer sub-block is processed by max pooling to obtain the feature vector of the current layer sub-block; Collect all sub-blocks of the current layer on all positive sample images and perform separate clustering; collect all sub-blocks of the current layer on all negative sample images and perform separate clustering to obtain the sub-cluster centers of the current layer; The images are scanned using the sub-cluster centers of the current layer, and the membership degree of each sub-block in the current layer is calculated as a feature value.
[0010] Preferably, calculating the responsivity of the target sub-block based on the color sub-cluster centers and the shape sub-cluster centers includes: Perform a full-image convolutional scan on the target image with a preset stride; At each sub-window location, if the target sub-block is a flat area, the average color value of the target sub-block is extracted, the responsivity is calculated using the color sub-cluster centers, and the responsivity corresponding to the shape sub-cluster centers is set to 0. If the target sub-block is a non-flat area, then the brightness of the target sub-block is normalized, the responsivity is calculated using the shape sub-cluster centers, and the responsivity corresponding to the color sub-cluster centers is set to 0.
[0011] Preferably, the step of quantifying the response into integers and calculating the distance between feature vectors using city distance includes: When calculating the responsivity, the range of the responsivity is linearly mapped from 0 to 1 to 0 to 255, and only integers are retained; The calculation is performed using the city distance on an integer basis, and the calculation of the city distance does not involve floating-point operations or multiplication operations; When calculating the responsivity of the target sub-block using the color sub-cluster center or the shape sub-cluster center, the city distance is calculated directly.
[0012] Preferably, generating the feature vector of the target image includes: In the last layer of a multi-layered sub-block, make the sub-window size cover the entire image; Based on the number of feature channels in the last layer and the preset step size, a final output feature vector of preset dimensions is obtained by concatenation, which serves as the feature vector of the target image.
[0013] Preferably, classifying the target image based on the feature vector includes: The final output feature vector is classified into two categories using the K-nearest neighbor algorithm, a neural network with only one hidden layer, a support vector machine, or a decision tree.
[0014] An interpretable image classification system according to the present invention includes: The region differentiation module is used to acquire sample sub-blocks of the sample image and, based on the difference between the maximum and minimum brightness values of the pixels within the sample sub-blocks, differentiate the sample sub-blocks into flat areas and non-flat areas. An independent clustering module is used to perform color clustering on sample sub-blocks belonging to the flat area and shape clustering on sample sub-blocks belonging to the non-flat area. During clustering, positive sample images and negative sample images are performed independently to obtain color sub-cluster centers and shape sub-cluster centers. The target region differentiation module is used to acquire target sub-blocks of the target image and, based on the difference between the maximum and minimum brightness values of the pixels within the target sub-blocks, differentiate the target sub-blocks into flat areas and non-flat areas. The responsiveness calculation module is used to calculate the responsiveness of the target sub-block based on the color sub-cluster centers and the shape sub-cluster centers; The feature generation module is used to quantize the response into integers and calculate the distance between feature vectors using city distance to generate feature vectors of the target image; A classification module is used to classify the target image based on the feature vector.
[0015] A computer device according to the present invention includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement an interpretable image classification method.
[0016] According to the present invention, a computer-readable storage medium is provided thereon storing a computer program that, when executed by a processor, implements an interpretable image classification method.
[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. This application distinguishes flat and non-flat areas based on the difference in pixel brightness within sub-blocks, and performs color clustering and shape clustering respectively. During clustering, positive and negative samples are processed independently, so that the feature extraction process no longer relies on black-box backpropagation updates, but instead constructs feature bases based on clear physical meaning and prior knowledge of categories. This mechanism makes each feature node of the model have a clear semantic interpretation, solves the problem of opaque internal mechanisms of deep neural networks, and enhances the discriminative power of features for specific categories by independent clustering of positive and negative samples, thereby improving the controllability and targeting of the model.
[0018] 2. This application eliminates a large number of floating-point operations and multiplication operations in the traditional deep learning inference process by quantizing the feature response into integers and using city distance to calculate the distance between feature vectors. City distance only involves addition and subtraction operations. Combined with the integer quantization strategy, the computational complexity and hardware resource consumption are greatly reduced while ensuring the accuracy of feature matching. This allows the image classification method to be efficiently deployed on mobile devices, embedded platforms and microcontrollers and other devices with limited computing power and power.
[0019] 3. This application constructs a hierarchical feature representation system through multi-layer sub-block max pooling and layer-by-layer independent clustering mechanism. The feature generation of each layer is based on the cluster center determined by the previous layer, rather than randomly initialized weights. This deterministic heuristic training process avoids the local optimum trap caused by randomness in end-to-end learning and the excessive dependence on hyperparameter tuning, thus lowering the research and development threshold. At the same time, due to the removal of redundant connections and parameters, the model file size is significantly reduced, making it easier to store and transmit. Attached Figure Description
[0020] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a schematic flowchart of an image classification method according to an embodiment of this application. Detailed Implementation
[0021] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0022] like Figure 1 As shown, this embodiment provides an image classification method. This method replaces the opaque backpropagation training process in traditional deep neural networks with deterministic feature extraction and clustering mechanisms, achieving complete transparency and interpretability of the model's internal mechanisms. In this embodiment, the image classification method is a binary classification algorithm, with two image classes: cars and flowers. The number of car images is 20,441, and the number of flower images is 17,512. It should be understood that the above dataset is merely illustrative, and the method of this application can be widely applied to other types of image classification tasks. The method mainly includes the following steps.
[0023] Step S1: Obtain sample sub-blocks of the sample image. Based on the difference between the maximum and minimum brightness values of pixels within the sample sub-blocks, classify the sample sub-blocks into flat and non-flat areas.
[0024] Specifically, before processing the sample image, it can be uniformly scaled to a preset size (e.g., 227×227 pixels) and converted to grayscale to standardize the input data. Then, a sample sub-block is cropped from the image using a sliding window of a preset size (e.g., 6×6 pixels). For each sample sub-block, the difference between the maximum and minimum brightness values of all pixels within it is calculated. When this difference is less than a preset threshold, the sample sub-block is determined to belong to a flat area; when the difference is greater than or equal to the preset threshold, the sample sub-block is determined to belong to a non-flat area. In this embodiment, the preset threshold is preferably set to 10. It should be understood that this threshold is not fixed and can be adjusted in practical applications according to the dynamic range of the image acquisition device or the scene noise level. For example, in low-contrast scenes, the threshold can be appropriately reduced to improve sensitivity. This physical judgment logic based on brightness differences can explicitly separate uniformly colored flat areas from structured areas containing edges and textures in the image, providing a clear physical basis for subsequent targeted extraction of color and shape features, avoiding abstract black-box feature learning.
[0025] Step S2: Perform color clustering on sample sub-blocks belonging to flat areas and shape clustering on sample sub-blocks belonging to non-flat areas. During clustering, positive and negative sample images are processed independently to obtain color sub-cluster centers and shape sub-cluster centers.
[0026] To address the different physical characteristics of flat and non-flat areas, this embodiment employs differentiated feature representation and clustering strategies: For sample sub-blocks belonging to flat areas, the average color value of the sub-block is calculated to obtain a color feature vector, and color clustering is performed based on this feature vector. Specifically, if the original image is a color image, the arithmetic mean of all pixels within the sub-block on the R, G, and B channels is calculated separately, and combined to form a 3D color feature vector. This vector can effectively represent the overall tonal information of the flat area. Subsequently, clustering algorithms such as K-Means are used to cluster the color feature vectors of all collected sample sub-blocks in the flat area, obtaining several color sub-cluster centers. For example, the number of clusters can be set to 100, thus obtaining 100 reference vectors representing different typical colors. During the clustering process, the average color value of each 6×6 sub-block is kept only as an integer value, and the city distance is selected as the distance metric to accelerate the clustering speed.
[0027] For sample sub-blocks belonging to non-flat areas, brightness normalization is performed on these sub-blocks to obtain shape feature vectors, and shape clustering is then performed based on these feature vectors. Specifically, to eliminate the interference of absolute illumination intensity on shape recognition, only the relative brightness distribution between pixels (i.e., shape structure information) is retained, and normalization is performed using the following linear mapping formula:
[0028] in, Indicates the first sub-block The original brightness value of each pixel. and These are the maximum and minimum brightness values within the sub-block, respectively. This is the normalized brightness value. (This refers to the brightness value of all pixels.) Arranging the shapes in spatial order yields a shape feature vector with a length equal to the total number of pixels in the sub-blocks (e.g., 36 dimensions). Similarly, the shape feature vectors of non-flat areas are clustered (e.g., the number of clusters is set to 100) to obtain shape sub-cluster centers. These centers represent typical local textures or edge patterns in the image. During shape clustering, the 36 normalized brightness values also retain only integer values and are clustered by city distance. The feature vectors corresponding to the cluster centers also retain only integer values, thus accelerating the clustering process.
[0029] It is particularly important to emphasize that in performing the aforementioned color and shape clustering, positive and negative sample images are processed independently. That is, flat area sub-blocks of positive samples participate in the same clustering process only with other flat area sub-blocks of positive samples, and the same applies to negative samples; the two are never mixed. This independent clustering mechanism is the key difference between this application and conventional unsupervised feature learning. Its technical effect is that by introducing prior knowledge of category labels, cluster centers are forced to converge towards the typical features of their respective categories, resulting in color and shape sub-cluster centers with strong category discriminative power. Compared to the feature space overlap and ambiguity caused by mixed positive and negative sample clustering, independent clustering ensures the semantic purity of the feature base, laying a solid foundation for subsequent high-precision classification.
[0030] To capture multi-scale information ranging from local details to global semantics, the clustering and feature extraction processes described above can be executed progressively layer by layer in a multi-layered sub-block architecture. Table 1 shows the specific parameter configurations for each layer of the network model in this embodiment.
[0031] Table 1 Parameters of each layer of the network model
[0032] As shown in Table 1, an 11-layer network model is constructed, and convolution is performed on the image at each layer to achieve stepwise feature extraction. Taking the second layer (6×6 sub-block layer) as an example, it performs convolution scanning on a 227×227 image with a stride of 2, and the final feature map size is 111×111×200 (100 color sub-cluster centers and 100 shape sub-cluster centers correspond to a feature map of 200 channels).
[0033] Within the current layer sub-block, multiple sub-blocks from the previous layer are included. For example, the third layer sub-block is 9×9 in size, covering four sub-blocks of the second layer (6×6) (i.e., arranged in a 2×2 pattern), with a preset interval of 3 pixels between adjacent second-layer sub-blocks. Max pooling is applied to the feature vectors corresponding to the previous layer sub-blocks to obtain the feature vector of the current layer sub-block. Specifically, for each channel of the previous layer feature map covered by the current layer sub-block, the maximum value among its multiple positional feature values is taken as the feature response of the current layer in that channel. Max pooling effectively filters out the most significant feature activation signals within the receptive field, enhancing the translation invariance and robustness of the features. Taking a 9×9 sub-block layer as an example, there are a total of 200 feature maps corresponding to the 6×6 sub-block layer. For each feature map, there are a total of 4 feature values corresponding to the current 9×9 sub-block. Max pooling is used to process these 4 feature values, and the maximum value among the 4 numbers is selected as the result after pooling. In this way, the 9×9 sub-block also obtains a feature vector of length 200.
[0034] After generating the feature vectors of the current layer sub-blocks, the independent clustering strategy described above is repeated: collect the current layer sub-blocks on all positive sample images and cluster them separately; collect the current layer sub-blocks on all negative sample images and cluster them separately, thus obtaining the sub-cluster centers of the current layer. As the layer deepens, the sub-block size gradually increases, and the number of cluster centers can also be increased accordingly to accommodate more complex semantic patterns. Using the sub-cluster centers of the current layer, each image is scanned, and the membership degree of each current layer sub-block is calculated as a feature value, thereby generating the feature map of that layer. Taking a 9×9 sub-block layer as an example, 9×9 sub-blocks on all positive sample images (each sub-block corresponds to a feature vector of length 200) are collected, and they are clustered, with the number of sub-clusters set to 300 (the larger the sub-block size, the more clusters). Similarly, during clustering in this layer, positive sample sub-blocks are clustered separately, and negative sample sub-blocks are clustered separately; the sub-blocks of the two classes are not mixed. The images are scanned using the 300 sub-cluster centers mentioned above, and the membership degree of each 9×9 sub-block is calculated as a feature value. The total number of feature maps obtained is 300.
[0035] Each subsequent layer is derived from the sub-blocks of the previous layer. By controlling the spacing of the sub-blocks in the previous layer (as shown in Table 1), the current layer's sub-blocks contain exactly four sub-blocks from the previous layer (i.e., a 2×2 arrangement). Similarly, all sub-blocks in the current layer are clustered, and the cluster centers are used for feature extraction. This process is repeated for each layer until the last layer (layer 11). Through this deterministic, layer-by-layer stacking feature extraction and clustering process, this embodiment constructs a deep feature pyramid with fully interpretable internal parameters and no error-driven backpropagation required. The feature generation of each layer directly originates from the explicit physical and statistical laws of the previous layer.
[0036] Step S3: Obtain the target sub-block of the target image. Based on the difference between the maximum and minimum brightness values of the pixels within the target sub-block, the target sub-block is divided into flat areas and non-flat areas.
[0037] Specifically, when processing the target image, the processing method is consistent with that for the sample image in step S1, using a sliding window of a preset size to extract target sub-blocks from the target image. For each target sub-block, the difference between the maximum and minimum brightness values of all pixels within it is calculated. When this difference is less than a preset threshold, the target sub-block is determined to belong to a flat area; when the difference is greater than or equal to the preset threshold, the target sub-block is determined to belong to a non-flat area. This determination logic is exactly the same as the region differentiation method for the sample sub-blocks in step S1, and will not be repeated here.
[0038] Step S4: Calculate the response of the target sub-block based on the color sub-cluster centers and shape sub-cluster centers.
[0039] After obtaining the color and shape sub-cluster centers obtained in the aforementioned training phase, a full-image convolutional scan of the target image is performed with a preset stride. For example, for the first layer of 6×6 sub-blocks, the convolutional stride can be set to 2, thereby generating a dense feature response map. The final feature map size is 111×111×200 (200 cluster centers corresponding to 200 channels of feature map). At each sub-window position, based on the region to which the target sub-block belongs determined in step S3, if the target sub-block is a flat area, the average color value of the target sub-block is extracted, the responsivity is calculated using the color sub-cluster centers, and the responsivity corresponding to the shape sub-cluster centers is set to 0; if the target sub-block is a non-flat area, the brightness of the target sub-block is normalized, the responsivity is calculated using the shape sub-cluster centers, and the responsivity corresponding to the color sub-cluster centers is set to 0. Therefore, the feature vector of each sub-block position is 200-dimensional. Here, the responsivity is the probability that the sub-block belongs to that sub-cluster. This channel zeroing logic has a dual technical effect: on the one hand, it achieves explicit decoupling of color features and shape features, preventing two types of features with vastly different physical meanings from interfering with each other in the same vector space, thus ensuring the purity of feature semantics; on the other hand, it makes the generated feature map highly sparse, significantly reducing the amount of invalid data in subsequent processing. It should be understood that although this embodiment uses a 6×6 sub-block as an example, the same conditional zeroing principle is followed in each layer of the multi-layer architecture, with only the sub-block size and step size adjusted according to the layer.
[0040] Step S5: Measure the response into integers and use city distance to calculate the distance between feature vectors to generate feature vectors of the target image.
[0041] To completely eliminate the high floating-point operation cost in traditional deep learning inference, this embodiment linearly maps the response range from 0 to 1 to 0 to 255 when calculating the response, retaining only the integers. The specific mapping formula is as follows: ,in For the original floating-point response, This is the quantized integer responsivity. Based on this, city distance is applied for feature matching calculation. The city distance referred to in this application, specifically the Manhattan distance, is the sum of the absolute values of the differences between two feature vectors across all dimensions. For example, for two N-dimensional integer feature vectors A and B, their city distance... The calculation process involves only subtraction, absolute value taking, and addition operations, completely excluding floating-point and multiplication operations. This design is the core mechanism for achieving low-power deployment in this application: compared to the square and square root operations required for Euclidean distance, city distance can be executed in parallel at high speed using only a basic ALU unit at the hardware level, enabling the algorithm to run in real time on computing-constrained platforms such as microcontrollers and DSPs, while maintaining sufficient feature discrimination.
[0042] In terms of clustering acceleration, the previous layer linearly maps the responsivity range from 0 to 1 to 0 to 255 when calculating the responsivity, retaining only integers. This ensures that the pooling results obtained after pooling in the current layer are also integers. During clustering, all feature vectors are converted to integers, allowing city distances to be applied on top of these integers for clustering, thus optimizing computation speed.
[0043] In the last layer of the multi-layer sub-blocks, the sub-window size covers the entire image, and based on the number of feature channels and the preset stride of the last layer, a final output feature vector of preset dimensions is concatenated as the feature vector of the target image. Specifically, the receptive field gradually expands with the progression of layers. Taking the 11-layer architecture described in Table 1 as an example, in the 11th layer (the last layer), the sub-window size reaches 211×211, almost covering the entire 227×227 image. Assuming that the number of feature channels in this layer is 500 and the convolution stride is 10, each channel generates 4 response values after scanning the entire image (corresponding to a 2×2 spatial grid). Concatenating these 500 channel response values in sequence yields a 2000-dimensional final output feature vector. This vector condenses complete information from local texture to global semantics, and because each layer undergoes independent clustering of positive and negative samples and integer compression, this 2000-dimensional vector has extremely high information density and class discriminative power, and can be directly used for classification without complex post-processing.
[0044] Step S6: Classify the target image based on the feature vector.
[0045] Since the front-end feature extraction process has already incorporated sufficient prior knowledge of the categories and achieved a highly structured representation, the back-end classifier does not need to possess complex learning capabilities. This embodiment employs the K-nearest neighbor algorithm, a neural network with only one hidden layer, a support vector machine, or a decision tree to perform binary classification on the final output feature vector. For example, when using the K-nearest neighbor algorithm, the city distance between the 2000-dimensional feature vector of the test sample and the sample vector of the training set is directly calculated, and the K nearest neighbors are selected for voting. This "feature extraction-heavy, classification-decision-light" architecture design makes the overall model file size extremely small, the training process completely deterministic and traceable, and completely avoids the inherent defects of end-to-end training of deep neural networks, such as parameter redundancy, black-box uninterpretability, and vulnerability to adversarial attacks.
[0046] In this embodiment, the classification error rate is calculated as follows: For example, in the current layer 2, 551 car images are classified as cars, and all remaining images are classified as flowers (flower images will not be misclassified). The number of misclassified images is 20441 - 551 = 19890, accounting for 97.3% of the car images. The classification error rate for flowers is 0. Therefore, the total classification error rate is (97.3% + 0) / 2 = 48.7%. As the layer deepens, the classification error rate gradually decreases, proving the effectiveness of multi-layer feature extraction.
[0047] Furthermore, this application also provides an embodiment of an image classification device. The device includes a region differentiation module, an independent clustering module, a target region differentiation module, a responsivity calculation module, a feature generation module, and a classification module. The system comprises the following modules: a region differentiation module for acquiring sample sub-blocks of the sample image, classifying them into flat and non-flat regions based on the difference between the maximum and minimum brightness values of pixels within each sub-block; an independent clustering module for performing color clustering on sample sub-blocks belonging to flat regions and shape clustering on sample sub-blocks belonging to non-flat regions, with positive and negative sample images clustered independently to obtain color and shape sub-cluster centers; a target region differentiation module for acquiring target sub-blocks of the target image, classifying them into flat and non-flat regions based on the difference between the maximum and minimum brightness values of pixels within each sub-block; a responsivity calculation module for calculating the responsivity of the target sub-blocks based on the color and shape sub-cluster centers, and performing channel zeroing during the calculation process; a feature generation module for converting the responsivity to integers and calculating the distance between feature vectors using city distance to generate feature vectors for the target image; and a classification module for classifying the target image based on the feature vectors. The functional implementation details of each module correspond one-to-one with steps S1 to S6 in the aforementioned method embodiment, and will not be repeated here. Through this modular architecture, the image classification method of this application can be flexibly packaged into a software SDK, firmware program or dedicated hardware accelerator, and is widely adapted to various smart terminals and edge computing devices.
[0048] This application also provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the various steps of the above-described image classification method. This computer device can be a server, a personal computer, an embedded device, or a mobile terminal, etc., and this application does not limit it to any of these.
[0049] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the various steps of the image classification method described above. The computer-readable storage medium can be a tangible, non-transitory storage medium, such as a disk, optical disk, flash memory, solid-state drive, etc., and this application does not limit it to this type.
[0050] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0051] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.
Claims
1. An interpretable image classification method, characterized in that, include: Step S1: Obtain sample sub-blocks of the sample image, and classify the sample sub-blocks into flat areas and non-flat areas based on the difference between the maximum and minimum brightness values of the pixels within the sample sub-blocks; Step S2: Perform color clustering on sample sub-blocks belonging to the flat area and shape clustering on sample sub-blocks belonging to the non-flat area. During clustering, positive and negative sample images are performed independently to obtain color sub-cluster centers and shape sub-cluster centers. Step S3: Obtain the target sub-block of the target image, and classify the target sub-block into flat area and non-flat area based on the difference between the maximum and minimum brightness values of the pixels within the target sub-block; Step S4: Calculate the responsivity of the target sub-block based on the color sub-cluster centers and the shape sub-cluster centers; Step S5: Measure the response into integers and use city distance to calculate the distance between feature vectors to generate the feature vector of the target image; Step S6: Classify the target image based on the feature vector.
2. The interpretable image classification method according to claim 1, characterized in that, The step of classifying the sample sub-block into flat and non-flat regions based on the difference between the maximum and minimum brightness values of pixels within the sample sub-block includes: When the difference between the maximum and minimum brightness values of all pixels in the sample sub-block is less than a preset threshold, the sample sub-block is determined as a flat area. When the difference is greater than or equal to the preset threshold, the sample sub-block is determined to be a non-flat area; The step of performing color clustering on sample sub-blocks belonging to the flat area includes: Calculate the average color value of the sample sub-blocks belonging to the flat area to obtain a color feature vector, and perform color clustering based on the color feature vector; The shape clustering of sample sub-blocks belonging to the non-flat area includes: Brightness normalization is performed on sample sub-blocks belonging to the non-flat area to obtain shape feature vectors, and shape clustering is performed based on the shape feature vectors.
3. The interpretable image classification method according to claim 2, characterized in that, The method also includes a feature extraction and clustering process for multi-layer sub-blocks: In the current layer sub-block, the current layer sub-block contains multiple previous layer sub-blocks; The feature vector corresponding to the previous layer sub-block is processed by max pooling to obtain the feature vector of the current layer sub-block; Collect all sub-blocks of the current layer on all positive sample images and perform separate clustering; collect all sub-blocks of the current layer on all negative sample images and perform separate clustering to obtain the sub-cluster centers of the current layer; The images are scanned using the sub-cluster centers of the current layer, and the membership degree of each sub-block in the current layer is calculated as a feature value.
4. The interpretable image classification method according to claim 1, characterized in that, The calculation of the responsivity of the target sub-block based on the color sub-cluster centers and the shape sub-cluster centers includes: Perform a full-image convolutional scan on the target image with a preset stride; At each sub-window location, if the target sub-block is a flat area, the average color value of the target sub-block is extracted, the responsivity is calculated using the color sub-cluster centers, and the responsivity corresponding to the shape sub-cluster centers is set to 0. If the target sub-block is a non-flat area, then the brightness of the target sub-block is normalized, the responsivity is calculated using the shape sub-cluster centers, and the responsivity corresponding to the color sub-cluster centers is set to 0.
5. The interpretable image classification method according to claim 4, characterized in that, The step of quantifying the response into an integer and calculating the distance between feature vectors using city distance includes: When calculating the responsivity, the range of the responsivity is linearly mapped from 0 to 1 to 0 to 255, and only integers are retained; The calculation is performed using the city distance on an integer basis, and the calculation of the city distance does not involve floating-point operations or multiplication operations; When calculating the responsivity of the target sub-block using the color sub-cluster center or the shape sub-cluster center, the city distance is calculated directly.
6. The interpretable image classification method according to claim 3, characterized in that, The generation of the feature vector of the target image includes: In the last layer of a multi-layered sub-block, make the sub-window size cover the entire image; Based on the number of feature channels in the last layer and the preset step size, a final output feature vector of preset dimensions is obtained by concatenation, which serves as the feature vector of the target image.
7. The interpretable image classification method according to claim 6, characterized in that, The classification of the target image based on the feature vector includes: The final output feature vector is classified into two categories using the K-nearest neighbor algorithm, a neural network with only one hidden layer, a support vector machine, or a decision tree.
8. An interpretable image classification system, characterized in that, include: The region differentiation module is used to acquire sample sub-blocks of the sample image and, based on the difference between the maximum and minimum brightness values of the pixels within the sample sub-blocks, differentiate the sample sub-blocks into flat areas and non-flat areas. An independent clustering module is used to perform color clustering on sample sub-blocks belonging to the flat area and shape clustering on sample sub-blocks belonging to the non-flat area. During clustering, positive sample images and negative sample images are performed independently to obtain color sub-cluster centers and shape sub-cluster centers. The target region differentiation module is used to acquire target sub-blocks of the target image and, based on the difference between the maximum and minimum brightness values of the pixels within the target sub-blocks, differentiate the target sub-blocks into flat areas and non-flat areas. The responsiveness calculation module is used to calculate the responsiveness of the target sub-block based on the color sub-cluster centers and the shape sub-cluster centers; The feature generation module is used to quantize the response into integers and calculate the distance between feature vectors using city distance to generate feature vectors of the target image; A classification module is used to classify the target image based on the feature vector.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the interpretable image classification method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the interpretable image classification method according to any one of claims 1 to 7.