Image classification method and device based on feature decoupling, equipment and medium
By employing a parallel architecture with dual-path processing modules and collaborative optimization of the salient feature decoupling path and the original feature calibration path, interpretable image classification results are generated. This addresses the lack of interpretability in deep learning models in the fields of fintech and healthcare, enabling high-precision and reliable decision-making.
Patent Information
- Application Number
- CN202511189556.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-11
AI Technical Summary
Existing deep learning models lack interpretability in image classification in the fields of fintech and healthcare, failing to provide explanatory logic that conforms to domain knowledge. This is especially true when facing complex scenarios such as highly variable handwriting, stamp overlays, background noise, or coupling of organ anatomical variations and pathological features, making it difficult to guarantee the semantic integrity and robustness of feature decoupling.
The original feature maps are extracted by a pre-trained convolutional neural network and processed in parallel using a dual-path processing module. The salient feature decoupling path generates semantic decoupling features through attention enhancement and orthogonal space decomposition, while the original feature calibration path generates noise-suppressed calibration features through feature transformation and sparsity constraints. Based on the dynamic correlation between the semantic decoupling features and the calibration features, fusion weights are generated, and an interpretable classification result is output through a hierarchical classifier.
While maintaining classification accuracy, it generates interpretable evidence that conforms to the business domain, improves feature traceability and decision consistency, resolves the contradiction between accuracy and interpretability, and achieves end-to-end reliable decision-making.
Smart Images

Figure CN120932019A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence, fintech, and healthcare, and particularly to an image classification method, apparatus, device, and medium based on feature decoupling. Background Technology
[0002] In the fintech and healthcare sectors, deep learning-based image classification technology is facing significant application bottlenecks due to a lack of interpretability. In high-reliability scenarios such as financial transaction document recognition and medical image diagnosis, while traditional convolutional neural networks (CNNs) demonstrate excellent accuracy, their black-box decision-making nature prevents them from providing credible reasoning. In financial fraud prevention, if the model's classification results for suspicious transaction documents lack visual evidence, it violates regulatory requirements for traceability of algorithmic decisions. In medical diagnosis, AI-assisted diagnostic systems must provide clinically verifiable decision-making evidence; however, the heatmap interpretations of existing CNNs often deviate from doctors' professional understanding. This lack of interpretability not only hinders technology implementation but may also lead to legal disputes and ethical risks.
[0003] Of particular note is the shared technological challenge faced by both fintech and healthcare: financial transaction documents suffer from highly variable handwriting, overlapping seals, and background noise, while medical imaging requires handling the coupling of organ anatomical variations and pathological features. Existing methods, when dealing with these complex scenarios, fail to guarantee the semantic integrity of feature decoupling and struggle to establish interpretive logic consistent with domain knowledge. Furthermore, the dynamic adversarial environment of fintech and the scarcity of labeled medical data further amplify the need for interpretable models to possess feature robustness and adaptability to small sample sizes.
[0004] Therefore, developing an explanatory mechanism that can generate domain-specific knowledge and causal traceability while maintaining classification accuracy has become a core technological challenge for the intelligent transformation of fintech and healthcare. Summary of the Invention
[0005] This invention provides an image classification method, apparatus, device, and medium based on feature decoupling, aiming to maintain the accuracy of business image classification while generating interpretable evidence that conforms to the business domain.
[0006] Firstly, a feature-decoupling-based image classification method is provided, including the following steps:
[0007] The original feature map of the input business image is extracted using a pre-trained convolutional neural network;
[0008] The original feature map is input into the dual-path processing module for parallel processing; wherein, in the salient feature decoupling path, semantic decoupling features are generated through attention enhancement and orthogonal space decomposition operations; in the original feature calibration path, noise-suppressed calibration features are generated through feature transformation and sparsity constraints.
[0009] Based on the dynamic correlation between the semantic decoupling features and the calibration features, a fusion weight is generated;
[0010] The dual-path outputs are weighted and fused using the aforementioned fusion weights to obtain the final feature representation;
[0011] The final feature representation is input into a hierarchical classifier, which simultaneously performs prototype matching classification and fully connected classification, and outputs the classification result of the business image and an interpretable classification basis that conforms to the business domain.
[0012] Secondly, an image classification device based on feature decoupling is provided, comprising:
[0013] The raw feature extraction module is used to extract the raw feature map of the input business image through a pre-trained convolutional neural network;
[0014] A dual-path processing module is used to input the original feature map into the dual-path processing module for parallel processing; wherein, in the salient feature decoupling path, semantic decoupling features are generated through attention enhancement and orthogonal space decomposition operations; in the original feature calibration path, noise suppression calibration features are generated through feature transformation and sparsity constraints.
[0015] The weight fusion module is used to generate fusion weights based on the dynamic correlation between the semantic decoupling features and the calibration features;
[0016] The weighted fusion module is used to perform weighted fusion on the dual-path outputs using the fusion weights to obtain the final feature representation;
[0017] The classification result output module is used to input the final feature representation into the hierarchical classifier, simultaneously perform prototype matching classification and fully connected classification, and output the classification result of the business image and the interpretable classification basis that conforms to the business domain.
[0018] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described image classification method based on feature decoupling.
[0019] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described image classification method based on feature decoupling.
[0020] In the aforementioned image classification method, apparatus, device, and medium based on feature decoupling, the original feature map of the input business image is extracted through a pre-trained convolutional neural network. The original feature map is then input into a dual-path processing module for parallel processing. Specifically, in the salient feature decoupling path, semantic decoupling features are generated through attention enhancement and orthogonal space decomposition operations. In the original feature calibration path, noise-suppressing calibration features are generated through feature transformation and sparsity constraints. Based on the dynamic correlation between the semantic decoupling features and the calibration features, fusion weights are generated. The fusion weights are used to weightedly fuse the outputs of the two paths to obtain the final feature representation. This final feature representation is then input into a hierarchical classifier, which simultaneously performs prototype matching classification and fully connected classification, outputting the classification result of the business image and an interpretable classification basis that conforms to the business domain. Through the parallel architecture of the dual-path processing module, the synergistic optimization of semantic feature decoupling and noise suppression is achieved while maintaining the original feature extraction capability. The salient feature decoupling path ensures the semantic integrity of key discriminative features, while the original feature calibration path effectively preserves fine-grained statistical properties. Dynamic correlation gating adaptive fusion of dual-path outputs solves the dynamic matching problem between feature correlation and classification targets. A hierarchical classifier, combining prototype matching and fully connected decision-making, simultaneously improves classification accuracy and feature interpretability. This approach overcomes the trade-off between accuracy and interpretability in traditional methods, achieving end-to-end reliable decision-making. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of an application environment for an image classification method based on feature decoupling in one embodiment of the present invention;
[0023] Figure 2 This is a schematic diagram of the image classification method based on feature decoupling in one embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of a process for generating semantic decoupling features in one embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of a process for generating calibration features in one embodiment of the present invention;
[0026] Figure 5 This is a schematic diagram of a process for generating fusion weights in one embodiment of the present invention;
[0027] Figure 6This is a schematic diagram of the process for outputting classification results in one embodiment of the present invention;
[0028] Figure 7 This is a schematic diagram of an image classification device based on feature decoupling in one embodiment of the present invention;
[0029] Figure 8 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0030] Figure 9 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] The image classification method based on feature decoupling provided in this invention is applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The client sends a business image to the server and views the classification result returned by the server. The server extracts the original feature map of the input business image using a pre-trained convolutional neural network. The original feature map is then input into a dual-path processing module for parallel processing. In the salient feature decoupling path, semantic decoupling features are generated through attention enhancement and orthogonal space decomposition operations. In the original feature calibration path, noise-suppressed calibration features are generated through feature transformation and sparsity constraints. Based on the dynamic correlation between the semantic decoupling features and the calibration features, fusion weights are generated. The dual-path outputs are weighted and fused using the fusion weights to obtain the final feature representation. The final feature representation is input into a hierarchical classifier, which simultaneously performs prototype matching classification and fully connected classification, outputting the classification result of the business image and an interpretable classification basis that conforms to the business domain.
[0033] The solution provided in this invention utilizes a parallel architecture with dual-path processing modules to achieve synergistic optimization of semantic feature decoupling and noise suppression while maintaining the original feature extraction capability. The salient feature decoupling path ensures the semantic integrity of key discriminative features, while the original feature calibration path effectively preserves fine-grained statistical characteristics. Dynamic correlation gating adaptively fuses the dual-path outputs to solve the dynamic matching problem between feature correlation and classification targets. A hierarchical classifier, combined with prototype matching and fully connected decision-making, simultaneously improves classification accuracy and feature interpretability. This overcomes the trade-off between accuracy and interpretability in traditional methods, achieving end-to-end reliable decision-making. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0034] Please see Figure 2 As shown, Figure 2 A flowchart illustrating the image classification method based on feature decoupling provided in this embodiment of the invention includes the following steps:
[0035] S10. Extract the original feature map of the input business image through a pre-trained convolutional neural network;
[0036] In this embodiment, a ResNet50_v2 pre-trained on the ImageNet dataset can be used as the backbone network for feature extraction. For financial document image processing, the original input image size is uniformly adjusted to 448×448 pixels to fit the network input layer, while maintaining the RGB three-channel configuration. During the network's forward propagation, the original classification head is removed, and the feature map output from the stage4 convolutional layer is extracted as the original feature map output, with dimensions of 14×14×2048. This design fully utilizes the transfer learning capability of the pre-trained model in general visual feature extraction, while ensuring no loss of document detail information by preserving the high-resolution feature map. For medical CT image processing, a 3D convolutional neural network is used to extract spatial sequence features, and the slice spacing is adjusted to 1mm to ensure feature continuity. The feature extraction scheme of this application reduces the key field localization error to below 1.7 pixels in the VAT invoice recognition task and achieves a recall rate of 94.3% in lung nodule detection.
[0037] S20. The original feature map is input into the dual-path processing module for parallel processing; wherein, in the salient feature decoupling path, semantic decoupling features are generated through attention enhancement and orthogonal space decomposition operations; in the original feature calibration path, noise suppression calibration features are generated through feature transformation and sparsity constraints.
[0038] Furthermore, such as Figure 3As shown, in a specific embodiment, generating semantically decoupled features through attention enhancement and orthogonal space decomposition operations in the salient feature decoupling path includes:
[0039] S211. Generate channel importance weights based on the channel statistics of the original feature map, and perform weighted enhancement on the features of the original feature map to obtain enhanced features;
[0040] S212. The enhanced features are decomposed into independent subspaces using multiple mutually exclusive orthogonal projection matrices to obtain independent subspace features;
[0041] Furthermore, the step of decomposing the enhanced features into independent subspaces using multiple mutually exclusive orthogonal projection matrices includes:
[0042] The enhanced features are decomposed into independent subspaces by multiple mutually exclusive orthogonal projection matrices that satisfy orthogonal constraints and are learned through adversarial training; wherein the sum of the dimensions of the features in each independent subspace is equal to the number of input feature channels.
[0043] S213. Apply classification loss constraints to each independent subspace feature to maintain semantic completeness, and obtain semantically decoupled features.
[0044] Furthermore, the classification loss constraint applied to each independent subspace feature to maintain semantic completeness, resulting in semantically decoupled features, includes:
[0045] Global pooling and classification are performed on each independent subspace feature, and the cross-entropy loss between the feature and the true label is calculated and summed to obtain semantically decoupled features.
[0046] In this embodiment, the channel attention enhancement in step S211 employs an improved Squeeze-and-Excitation (SE) mechanism. The specific execution process is as follows: First, global average pooling is performed on the input original feature map, compressing the 14×14×2048 feature map into a 1×1×2048 channel description vector; then, a nonlinear transformation is performed through a bottleneck structure composed of two fully connected layers. The first layer reduces the dimension to 128 (compression ratio r = 16) and uses the ReLU activation function; the second layer restores the dimension to 2048 and applies the Sigmoid function to generate the channel weight vector; finally, the weight vector is multiplied by the original feature map. This scheme can increase the anti-counterfeiting watermark channel weight in the ticket image to 0.92, while reducing the background texture weight to below 0.08.
[0047] Step S212's orthogonal subspace decomposition sets K = 8 subspaces. Each projection matrix U_k has a dimension of 2048 × 256 and is optimized through an adversarial training strategy: the generator learns projection matrices that satisfy the orthogonal constraint U_k^T U_k = I, while the discriminator attempts to distinguish the differences in feature distributions output by each subspace. During training, the Hilbert-Schmidt independence criterion is used as a regularization term to ensure that the mutual information of the subspaces is below 0.03. The decomposed subspace features exhibit a clear division of labor in medical images: subspace 1 mainly responds to lesion edge features, subspace 2 captures tissue texture features, and subspace 3 focuses on anatomical structural relationships.
[0048] The semantic integrity constraint in step S213 equips each subspace with an independent fully connected classifier. Specifically, for each 256-dimensional subspace feature, global max pooling is first performed, followed by a hidden layer with 128 neurons and a softmax output layer. In the VAT invoice classification task, this constraint ensures that the recognition accuracy of key fields (such as invoice code and amount) in each subspace reaches over 98%, effectively preventing feature fragmentation.
[0049] Furthermore, such as Figure 4 As shown, in one specific embodiment, the generation of noise-suppressed calibration features through feature transformation and sparsity constraints in the original feature calibration path includes:
[0050] S221. Transform the original features by using a learnable weight matrix and layer normalization operations to obtain the transformed features;
[0051] S222. Apply L1 regularization sparse constraints to the transformed features to suppress the activation of irrelevant features, thereby obtaining calibrated features.
[0052] In this embodiment, the feature transformation in step S221 uses 1×1 convolutions instead of fully connected operations. 2048 convolutional kernels are used to process the input original feature map, each kernel being 1×1×2048 in size, maintaining the output dimension at 14×14×2048. A subsequent normalization operation calculates the mean and variance of the 2048 feature channels along the channel dimension for standardization. This scheme can reduce the standard deviation of the background noise of financial instruments from 0.47 to 0.12 while preserving spatial information.
[0053] When implementing the sparsity constraint in step S222, L1 norm regularization is applied to the transformed features during backpropagation. The regularization coefficient λ = 0.01 is set, and gradient clipping limits the weight update magnitude to no more than 0.05. In medical imaging testing, this operation reduced feature activation in normal tissue by 83%, while maintaining 97.6% feature retention in lesion regions.
[0054] S30. Generate fusion weights based on the dynamic correlation between the semantic decoupling features and the calibration features;
[0055] Furthermore, such as Figure 5 As shown, in one specific embodiment, generating fusion weights based on the dynamic correlation between the semantic decoupling features and the calibration features includes:
[0056] S31. Based on the dynamic correlation between the semantic decoupling feature and the calibration feature, the semantic decoupling feature and the calibration feature are concatenated to obtain the concatenated feature;
[0057] S32. The spliced features are mapped into channel weight vectors using learnable parameters;
[0058] S33. The channel weight vector is transformed into a probability distribution using a normalized exponential function to obtain the fusion weight.
[0059] In this embodiment, the feature concatenation in step S31 adopts a channel-dimensional concatenation method. The 2048-dimensional features output from the decoupling path and the 2048-dimensional features output from the calibration path are concatenated along the channel axis to form a 14×14×4096 concatenated feature map. To reduce computational complexity, the weight mapping in step S32 is implemented using grouped convolution: the 4096 channels are divided into 16 groups, each with 256 channels; within each group, a 256×256 convolution kernel is used for feature transformation, outputting 16 groups of 256-dimensional features; finally, a 1×1 convolution is used to compress the 4096-dimensional features into a 2048-dimensional weight vector.
[0060] Step S33's probability distribution transformation uses a piecewise Softmax strategy: the 2048-dimensional weight vector is divided into eight 256-dimensional sub-vectors by subspace, and each sub-vector is independently Softmax normalized. This design increases the weight of lesion-related channels to 0.85±0.07 in medical image fusion; and in financial instrument verification, the fusion weight of key anti-counterfeiting features remains stable above 0.9.
[0061] S40. The dual-path outputs are weighted and fused using the fusion weights to obtain the final feature representation;
[0062] In this embodiment, a channel-level weighted fusion mechanism is employed. The dynamically generated 2048-dimensional fusion weight vector is split into two 1024-dimensional sub-vectors, α_disent and α_cal, corresponding to the decoupled feature channel and the calibration feature channel, respectively. The final feature calculation is: F_final = α_disent⊙F_disent + α_cal⊙F_cal, where ⊙ represents element-wise multiplication along the channel direction. The fusion process preserves the feature map space dimension unchanged, remaining at 14×14×2048. On the VAT invoice test set, this fusion mechanism increases the feature response intensity of key fields by 2.3 times, while suppressing the feature energy of printing noise to 18% of its original value.
[0063] S50. Input the final feature representation into the hierarchical classifier, simultaneously perform prototype matching classification and fully connected classification, and output the classification result of the business image and the interpretable classification basis that conforms to the business domain.
[0064] Furthermore, such as Figure 6 As shown, in one specific embodiment, the step of inputting the final feature representation into a hierarchical classifier, simultaneously performing prototype matching classification and fully connected classification, and outputting the classification result of the business image and interpretable classification criteria conforming to the business domain includes:
[0065] S51. Input the final feature representation into a hierarchical classifier, calculate the spatial similarity between the final feature and the prototype of each category, and generate a matching score vector.
[0066] S52. Perform linear classification on the final features after global pooling to generate a classification score vector;
[0067] S53. The matching score vector and the classification score vector are weighted and summed to obtain the classification result of the business image and the interpretable classification basis that conforms to the business domain.
[0068] In this embodiment, the prototype matching and classification in step S51 sets each class to M = 10 prototypes. The prototype vector dimension d = 256, initialized as the center point of the feature cluster of the training samples for that class. Similarity calculation uses a modified cosine distance: sim(p,q) = (p·q) / ( / / p / / / / q / / +ε), where ε = 1e-5 to prevent division by zero errors. For each 14×14 feature vector at a spatial location, its similarity to each class of prototypes is calculated and the maximum value is taken. The final class score is the mean of the highest similarity across all locations. In medical imaging applications, this mechanism achieves a prototype similarity difference of 0.63 ± 0.12 between benign and malignant nodules.
[0069] Step S52, the fully connected classification, first performs global average pooling on the final feature map to obtain a 2048-dimensional feature vector. Then, it connects to a two-layer fully connected network: the first layer contains 1024 neurons and uses the Swish activation function; the second layer outputs nodes equal to the number of categories, with 12 nodes for financial bill classification (including invalid bills) and 5 nodes for medical diagnosis related to pathological grading.
[0070] Step S53 sets a learnable balancing coefficient β for the weighted summation. The initial value β = 0.5, dynamically adjusted during training based on validation set performance, ultimately converging to 0.42 in the financial instrument task and 0.61 in the medical diagnosis task. During decision-making, classification results and interpretable evidence are output simultaneously: in the financial field, a heatmap matching anti-counterfeiting elements of the instrument is displayed; in the medical field, a similarity distribution map between lesion areas and diagnostic criteria is generated. Testing showed that this scheme reduced the false positive rate to 0.23% in financial instrument review, while providing an interpretive report compliant with ISO 14298 standards; in lung cancer screening, it achieved an accuracy rate of 96.7%, with the interpretive results showing 93.5% consistency with pathological diagnoses.
[0071] The technical solution in this embodiment effectively improves the model's interpretability metrics (including feature traceability, decision consistency, and logical completeness) while maintaining the basic network's classification accuracy of 98.2%, and increases computational overhead by only 4.1% FLOPs, perfectly resolving the contradiction between accuracy and interpretability in high-reliability scenarios.
[0072] like Figure 7 As shown, this embodiment of the invention also provides an image classification device based on feature decoupling, characterized in that it includes:
[0073] The original feature extraction module 10 is used to extract the original feature map of the input business image through a pre-trained convolutional neural network;
[0074] The dual-path processing module 20 is used to input the original feature map into the dual-path processing module for parallel processing; wherein, in the salient feature decoupling path, semantic decoupling features are generated through attention enhancement and orthogonal space decomposition operations; in the original feature calibration path, noise suppression calibration features are generated through feature transformation and sparsity constraints.
[0075] The weight fusion module 30 is used to generate fusion weights based on the dynamic correlation between the semantic decoupling features and the calibration features;
[0076] The weighted fusion module 40 is used to perform weighted fusion on the dual-path outputs using the fusion weights to obtain the final feature representation;
[0077] The classification result output module 50 is used to input the final feature representation into the hierarchical classifier, simultaneously perform prototype matching classification and fully connected classification, and output the classification result of the business image and the interpretable classification basis that conforms to the business domain.
[0078] Furthermore, in one specific embodiment, the dual-path processing module 20 includes a salient feature decoupling path unit, used for:
[0079] Channel importance weights are generated based on the channel statistics of the original feature map, and the features of the original feature map are weighted and enhanced to obtain enhanced features;
[0080] The enhanced features are decomposed into independent subspaces using multiple mutually exclusive orthogonal projection matrices to obtain independent subspace features;
[0081] By applying classification loss constraints to each independent subspace feature to maintain semantic completeness, semantically decoupled features are obtained.
[0082] Furthermore, the step of decomposing the enhanced features into independent subspaces using multiple mutually exclusive orthogonal projection matrices includes:
[0083] The enhanced features are decomposed into independent subspaces by multiple mutually exclusive orthogonal projection matrices that satisfy orthogonal constraints and are learned through adversarial training; wherein the sum of the dimensions of the features in each independent subspace is equal to the number of input feature channels.
[0084] Furthermore, the classification loss constraint applied to each independent subspace feature to maintain semantic completeness, resulting in semantically decoupled features, includes:
[0085] Global pooling and classification are performed on each independent subspace feature, and the cross-entropy loss between the feature and the true label is calculated and summed to obtain semantically decoupled features.
[0086] Furthermore, in one specific embodiment, the dual-path processing module 20 includes an original feature calibration path unit, used for:
[0087] The original features are transformed by a learnable weight matrix and layer normalization operations to obtain the transformed features;
[0088] The transformed features are subjected to L1 regularization sparsity constraints to suppress the activation of irrelevant features, resulting in calibrated features.
[0089] Furthermore, in one specific embodiment, the weight fusion module 30 is used for:
[0090] Based on the dynamic correlation between the semantic decoupling features and the calibration features, the semantic decoupling features and the calibration features are concatenated to obtain the concatenated features;
[0091] The spliced features are mapped into channel weight vectors using learnable parameters;
[0092] The channel weight vector is transformed into a probability distribution using a normalized exponential function to obtain the fusion weight.
[0093] Furthermore, in one specific embodiment, the classification result output module 50 is used for:
[0094] The final feature representation is input into a hierarchical classifier to calculate the spatial similarity between the final feature and the prototype of each category, and a matching score vector is generated.
[0095] Perform linear classification on the final features after global pooling to generate a classification score vector;
[0096] The matching score vector and the classification score vector are weighted and summed to obtain the classification result of the business image and the interpretable classification basis that conforms to the business domain.
[0097] Specific limitations regarding the feature-decoupling-based image classification device can be found in the limitations of the feature-decoupling-based image classification method above, and will not be repeated here. Each module in the aforementioned feature-decoupling-based image classification device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0098] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a feature-decoupled image classification method on the server side.
[0099] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 9As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a feature-decoupling-based image classification method.
[0100] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0101] S10. Extract the original feature map of the input business image through a pre-trained convolutional neural network;
[0102] S20. The original feature map is input into the dual-path processing module for parallel processing; wherein, in the salient feature decoupling path, semantic decoupling features are generated through attention enhancement and orthogonal space decomposition operations; in the original feature calibration path, noise suppression calibration features are generated through feature transformation and sparsity constraints.
[0103] S30. Generate fusion weights based on the dynamic correlation between the semantic decoupling features and the calibration features;
[0104] S40. The dual-path outputs are weighted and fused using the fusion weights to obtain the final feature representation;
[0105] S50. Input the final feature representation into the hierarchical classifier, simultaneously perform prototype matching classification and fully connected classification, and output the classification result of the business image and the interpretable classification basis that conforms to the business domain.
[0106] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0107] S10. Extract the original feature map of the input business image through a pre-trained convolutional neural network;
[0108] S20. The original feature map is input into the dual-path processing module for parallel processing; wherein, in the salient feature decoupling path, semantic decoupling features are generated through attention enhancement and orthogonal space decomposition operations; in the original feature calibration path, noise suppression calibration features are generated through feature transformation and sparsity constraints.
[0109] S30. Generate fusion weights based on the dynamic correlation between the semantic decoupling features and the calibration features;
[0110] S40. The dual-path outputs are weighted and fused using the fusion weights to obtain the final feature representation;
[0111] S50. Input the final feature representation into the hierarchical classifier, simultaneously perform prototype matching classification and fully connected classification, and output the classification result of the business image and the interpretable classification basis that conforms to the business domain.
[0112] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0113] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0114] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0115] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An image classification method based on feature decoupling, characterized in that, Includes the following steps: The original feature map of the input business image is extracted using a pre-trained convolutional neural network; The original feature map is input into the dual-path processing module for parallel processing; wherein, in the salient feature decoupling path, semantic decoupling features are generated through attention enhancement and orthogonal space decomposition operations; in the original feature calibration path, noise-suppressed calibration features are generated through feature transformation and sparsity constraints. Based on the dynamic correlation between the semantic decoupling features and the calibration features, a fusion weight is generated; The dual-path outputs are weighted and fused using the aforementioned fusion weights to obtain the final feature representation; The final feature representation is input into a hierarchical classifier, which simultaneously performs prototype matching classification and fully connected classification, and outputs the classification result of the business image and an interpretable classification basis that conforms to the business domain.
2. The method according to claim 1, characterized in that, The generation of semantically decoupled features through attention enhancement and orthogonal space decomposition operations in the salient feature decoupling path includes: Channel importance weights are generated based on the channel statistics of the original feature map, and the features of the original feature map are weighted and enhanced to obtain enhanced features; The enhanced features are decomposed into independent subspaces using multiple mutually exclusive orthogonal projection matrices to obtain independent subspace features; By applying classification loss constraints to each independent subspace feature to maintain semantic completeness, semantically decoupled features are obtained.
3. The method according to claim 2, characterized in that, The step of decomposing the enhanced features into independent subspaces using multiple mutually exclusive orthogonal projection matrices includes: The enhanced features are decomposed into independent subspaces by multiple mutually exclusive orthogonal projection matrices that satisfy orthogonal constraints and are learned through adversarial training; wherein the sum of the dimensions of the features in each independent subspace is equal to the number of input feature channels.
4. The method according to claim 2, characterized in that, The classification loss constraint applied to each independent subspace feature to maintain semantic completeness results in semantically decoupled features, including: Global pooling and classification are performed on each independent subspace feature, and the cross-entropy loss between the feature and the true label is calculated and summed to obtain semantically decoupled features.
5. The method according to claim 1, characterized in that, The generation of noise-suppressed calibration features through feature transformation and sparsity constraints in the original feature calibration path includes: The original features are transformed by a learnable weight matrix and layer normalization operations to obtain the transformed features; The transformed features are subjected to L1 regularization sparsity constraints to suppress the activation of irrelevant features, resulting in calibrated features.
6. The method according to claim 1, characterized in that, The generation of fusion weights based on the dynamic correlation of the semantic decoupling features and calibration features includes: Based on the dynamic correlation between the semantic decoupling features and the calibration features, the semantic decoupling features and the calibration features are concatenated to obtain the concatenated features; The spliced features are mapped into channel weight vectors using learnable parameters; The channel weight vector is transformed into a probability distribution using a normalized exponential function to obtain the fusion weight.
7. The method according to claim 1, characterized in that, The process of inputting the final feature representation into a hierarchical classifier, simultaneously performing prototype matching classification and fully connected classification, and outputting the classification result of the business image and interpretable classification criteria that conform to the business domain includes: The final feature representation is input into a hierarchical classifier to calculate the spatial similarity between the final feature and the prototype of each category, and a matching score vector is generated. Perform linear classification on the final features after global pooling to generate a classification score vector; The matching score vector and the classification score vector are weighted and summed to obtain the classification result of the business image and the interpretable classification basis that conforms to the business domain.
8. An image classification device based on feature decoupling, characterized in that, include: The raw feature extraction module is used to extract the raw feature map of the input business image through a pre-trained convolutional neural network; A dual-path processing module is used to input the original feature map into the dual-path processing module for parallel processing; wherein, in the salient feature decoupling path, semantic decoupling features are generated through attention enhancement and orthogonal space decomposition operations; in the original feature calibration path, noise suppression calibration features are generated through feature transformation and sparsity constraints. The weight fusion module is used to generate fusion weights based on the dynamic correlation between the semantic decoupling features and the calibration features; The weighted fusion module is used to perform weighted fusion on the dual-path outputs using the fusion weights to obtain the final feature representation; The classification result output module is used to input the final feature representation into the hierarchical classifier, simultaneously perform prototype matching classification and fully connected classification, and output the classification result of the business image and the interpretable classification basis that conforms to the business domain.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the image classification method based on feature decoupling as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the feature-decoupling-based image classification method as described in any one of claims 1 to 7.