An image classification method and related equipment for cross-domain transfer learning

The heterogeneous feature extraction network and dynamic domain similarity measurement module generate a cross-domain migration weight matrix, combined with the deformable feature pyramid for spatial transformation and channel reorganization, solving the problem of fine-grained distribution differences between domains in cross-domain transfer learning, and improving feature matching accuracy and robustness.

CN120182726BActive Publication Date: 2025-08-08BYZORO NETWORK LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510648948.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-08
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing cross-domain transfer learning methods are difficult to effectively capture the fine-grained distribution differences between domains when the distributions of training data and test data are significant, resulting in a decrease in classification accuracy.

Method used

The multi-level semantic features of the source domain and the target domain are extracted through heterogeneous feature extraction network, combined with the dynamic domain similarity measurement module to generate a cross-domain migration weight matrix, and use the deformable feature pyramid to realize spatial transformation and channel reorganization, and finally build a target domain adaptive feature representation.

Benefits of technology

It improves feature matching accuracy and classification robustness in cross-domain scenarios, enhances adaptability and robustness to complex domain differences, and reduces computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182726B_ABST
    Figure CN120182726B_ABST
Patent Text Reader

Abstract

The present application discloses an image classification method and related equipment for cross-domain transfer learning, which relates to the field of image classification technology. The method includes: obtaining a source domain image set and a target domain image set; extracting first multi-level semantic features of the source domain image set and second multi-level semantic features of the target domain image set based on a heterogeneous feature extraction network; generating a cross-domain transfer weight matrix based on the first multi-level semantic features and the second multi-level semantic features through a dynamic domain similarity measurement module; performing spatial transformation on the first multi-level semantic features based on a deformable feature pyramid to generate a transfer feature map; performing channel reorganization on the transfer feature map based on the cross-domain transfer weight matrix to construct a target domain adaptation feature representation; and outputting the classification result of the target domain image through a target domain classifier based on the target domain adaptation feature representation. The present application improves the feature matching accuracy and classification robustness in cross-domain scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image classification technology, and in particular to an image classification method and related equipment for cross-domain transfer learning. Background Art

[0002] With the rapid development of computer vision technology, image classification tasks have been widely used in multiple fields. However, practical applications often face the problem of significant differences in the distribution of training data and test data, such as differences in lighting conditions, shooting angles, or object morphology, which leads to a sharp decline in the performance of traditional models in cross-domain scenarios. Existing cross-domain transfer learning methods are mostly based on adversarial training or feature alignment strategies, but they generally suffer from the problem of coarse granularity in feature space matching, making it difficult to effectively capture fine-grained distribution differences between domains, affecting classification accuracy. Therefore, there is an urgent need for an image classification method based on cross-domain transfer learning to address the above-mentioned technical issues. Summary of the Invention

[0003] The Summary of the Invention introduces a series of simplified concepts that will be further described in the Detailed Description of the Invention. The Summary of the Invention of this application is not intended to limit the key features and essential technical features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.

[0004] In a first aspect, the present application provides an image classification method for cross-domain transfer learning, the method comprising:

[0005] Obtain a source domain image set and a target domain image set;

[0006] Based on the heterogeneous feature extraction network, the first multi-level semantic features of the source domain image set and the second multi-level semantic features of the target domain image set are extracted;

[0007] Based on the first multi-level semantic features and the second multi-level semantic features, a cross-domain transfer weight matrix is generated through a dynamic domain similarity measurement module;

[0008] Based on the deformable feature pyramid, the first multi-level semantic features are spatially transformed to generate a migration feature map;

[0009] Based on the cross-domain transfer weight matrix, the transfer feature map is channel-reorganized to construct the target domain adaptation feature representation;

[0010] Based on the target domain adaptation feature representation, the classification result of the target domain image is output through the target domain classifier.

[0011] In some embodiments, extracting first multi-level semantic features of a source domain image set based on a heterogeneous feature extraction network includes:

[0012] Based on the source domain image set, shallow texture features are extracted by dilating convolutional layers with a preset number in the first convolutional branch of the heterogeneous feature extraction network, wherein the dilation rate of the convolution kernel of the first convolutional branch is negatively correlated with the image resolution;

[0013] Based on the source domain image set, deep semantic features are extracted through the cascaded atrous spatial pyramid pooling module in the second convolutional branch of the heterogeneous feature extraction network, where the cascaded atrous spatial pyramid pooling module includes multiple sets of parallel convolutional layers with different atrous rates;

[0014] Based on shallow texture features and deep semantic features, a first multi-level semantic feature is generated through a cross-resolution stitching operation, wherein the cross-resolution stitching operation includes feature map size alignment and channel dimension superposition.

[0015] In some embodiments, generating a cross-domain transfer weight matrix based on the first multi-level semantic features and the second multi-level semantic features by a dynamic domain similarity measurement module includes:

[0016] Based on the first multi-level semantic features, the cross-level distribution difference measurement is performed through the maximum mean difference calculation module to generate the source domain feature distribution vector;

[0017] Based on the second multi-level semantic features, probability density modeling is performed through the kernel density estimation module to generate the target domain feature distribution vector;

[0018] Based on the source domain feature distribution vector and the target domain feature distribution vector, interactive similarity matching is performed through a dual-channel attention mechanism to generate global attention weights and local attention weights. The global attention weight is generated by calculating the global correlation between the source domain feature distribution vector and the target domain feature distribution vector based on cosine similarity through the first channel of the dual-channel attention mechanism; the local attention weight is generated by calculating the local adaptability of the source domain feature distribution vector and the target domain feature distribution vector based on element-by-element difference through the second channel of the dual-channel attention mechanism;

[0019] Based on the weighted fusion results of global attention weights and local attention weights, a cross-domain transfer weight matrix is generated.

[0020] In some embodiments, performing spatial transformation on the first multi-level semantic features based on the deformable feature pyramid to generate a migration feature map includes:

[0021] Determine the deformation parameters of the deformable convolution kernel at each level in the deformable feature pyramid based on the local gradient features of the target domain image set, where the local gradient features are extracted through the shallow feature map of the second multi-level semantic features;

[0022] Based on the deformation parameters, the spatial offset of the deformable convolution kernel is adjusted to generate a deformable convolution kernel that is adaptive to the spatial distribution of the target domain;

[0023] Based on the deformable convolution kernel, bilinear interpolation is performed on the first multi-level semantic features to generate an intermediate feature map after position correction;

[0024] Based on the multi-level structure of the deformable feature pyramid, cross-level feature fusion is performed on the intermediate feature maps to generate a migration feature map that is aligned with the spatial distribution of the target domain.

[0025] In some embodiments, based on the cross-domain migration weight matrix, channel reorganization is performed on the migration feature map to construct a target domain adaptation feature representation, including:

[0026] Based on the channel dimension of the migration feature map, the channel attention weight is calculated through the cross-domain migration weight matrix to generate the channel attention vector;

[0027] Based on the channel attention vector, each channel feature of the migration feature map is weighted channel by channel to generate a weighted feature map;

[0028] Based on the weighted feature map, cross-level channel splicing is performed through a preset feature fusion strategy to generate a spliced feature map;

[0029] Based on the channel dimension of the spliced feature map, the channel dimension is compressed by a feature compression module with a preset dimensionality reduction ratio to generate a target domain adapted feature representation, where the channel dimension of the target domain adapted feature representation matches the input dimension of the target domain classifier.

[0030] In some embodiments, outputting a classification result of a target domain image by a target domain classifier based on the target domain adapted feature representation includes:

[0031] Based on the channel dimension of the target domain adaptation feature representation, the feature dimension reduction mapping is performed through the fully connected layer in the target domain classifier to generate a classification feature vector;

[0032] Based on the classification feature vector, a nonlinear transformation is performed through a preset activation function to generate a category probability distribution;

[0033] The classification result of the target domain image is determined based on the category label corresponding to the maximum probability value in the category probability distribution.

[0034] In some embodiments, determining deformation parameters of a deformable convolution kernel at each level in a deformable feature pyramid based on local gradient features of a target domain image set includes:

[0035] Based on the shallow feature map of the second multi-level semantic feature, a directional convolution operation is performed through a preset gradient operator to generate a horizontal gradient component and a vertical gradient component;

[0036] Based on the vector synthesis results of the transverse gradient component and the longitudinal gradient component, the spatial gradient direction distribution of the local gradient feature is determined;

[0037] Based on the statistical histogram of spatial gradient direction distribution, the dominant gradient direction corresponding to each level is extracted through clustering algorithm;

[0038] Based on the angle between the dominant gradient direction and the preset reference direction, the horizontal deformation offset and vertical deformation offset of the deformable convolution kernel of the corresponding level in the deformable feature pyramid are calculated;

[0039] Based on the horizontal deformation offset and the vertical deformation offset, the deformation parameters of the deformable convolution kernel are determined.

[0040] In a second aspect, the present application proposes an image classification device for cross-domain transfer learning, comprising:

[0041] An image data acquisition unit, configured to acquire a source domain image set and a target domain image set;

[0042] A semantic feature extraction unit, which extracts first multi-level semantic features of a source domain image set and second multi-level semantic features of a target domain image set based on a heterogeneous feature extraction network;

[0043] A weight matrix generating unit, which generates a cross-domain migration weight matrix based on the first multi-level semantic features and the second multi-level semantic features through a dynamic domain similarity measurement module;

[0044] A feature map generation unit, which performs spatial transformation on the first multi-level semantic features based on a deformable feature pyramid to generate a migration feature map;

[0045] The feature representation construction unit reorganizes the channels of the migrated feature map based on the cross-domain migration weight matrix to construct the target domain adaptation feature representation;

[0046] The image classification output unit outputs the classification results of the target domain image through the target domain classifier based on the target domain adaptation feature representation.

[0047] In a third aspect, an electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to execute the computer program stored in the memory to implement the steps of the image classification method of cross-domain transfer learning of any one of the first aspects.

[0048] In a fourth aspect, the present application proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image classification method of cross-domain transfer learning of any one of the first aspects.

[0049] In summary, this application extracts multi-level semantic features of the source domain and the target domain through a heterogeneous feature extraction network, combines the dynamic domain similarity measurement module to generate a cross-domain migration weight matrix, and uses a deformable feature pyramid to achieve spatial transformation and channel reorganization, and finally constructs a target domain adaptation feature representation. This application can accurately quantify the hierarchical differences in cross-domain feature distribution, adaptively adjust feature migration weights, and effectively align local structural features of the target domain through nonlinear spatial transformation. Compared with the existing technology, it improves the feature matching accuracy and classification robustness in cross-domain scenarios, while reducing computational overhead and enhancing the adaptability and robustness to complex domain differences.

[0050] The image classification method of cross-domain transfer learning proposed in this application, and other advantages, objectives and features of this application will be reflected in part through the following description, and in part will also be understood by technical personnel in this field through research and practice of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present description. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0052] Figure 1 A schematic diagram of the image classification process for cross-domain transfer learning provided in an embodiment of the present application;

[0053] Figure 2 A schematic diagram of the structure of an image classification device for cross-domain transfer learning provided in an embodiment of the present application;

[0054] Figure 3 A structural diagram of an electronic device for image classification using cross-domain transfer learning provided in an embodiment of the present application. DETAILED DESCRIPTION

[0055] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices. The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments.

[0056] See also Figure 1 , which is a flow chart of an image classification method for cross-domain transfer learning provided in an embodiment of the present application, which may specifically include:

[0057] S110, obtaining a source domain image set and a target domain image set;

[0058] For example, the acquisition of source domain image sets and target domain image sets is a prerequisite for cross-domain transfer learning. The source domain usually refers to a training data set with rich annotated information, and its data distribution is significantly different from that of the target domain, such as inter-domain offset caused by different acquisition equipment, environmental conditions or target morphology. The target domain image set represents the data to be classified in actual application scenarios, and its annotations may be limited or completely missing. The acquisition of both must fully consider the actual task requirements. For example, in medical image classification, the source domain may be high-quality images taken by standard laboratory equipment, while the target domain may be low-resolution or noisy images in a clinical environment. By clearly defining the differences between the source domain and the target domain, data support is provided for subsequent cross-domain migration.

[0059] The acquisition process must ensure the diversity and representativeness of the source and target domain image sets to cover the complex scenarios that may exist in the target domain. For example, in cross-camera scenarios, it is necessary to collect target domain samples from different angles and lighting conditions to simulate the distribution changes in actual deployments. The goal of this step is to provide the model with sufficient data comparison foundation, thereby effectively reducing the impact of inter-domain distribution differences on classification performance in subsequent feature alignment and transfer, and laying the data foundation for the implementation of cross-domain transfer learning methods.

[0060] S120, extracting first multi-level semantic features of the source domain image set and second multi-level semantic features of the target domain image set based on the heterogeneous feature extraction network;

[0061] For example, the heterogeneous feature extraction network captures the multi-level semantic information of the source domain and the target domain respectively through differentiated branch structures. For the source domain image set, the network works together through the dual branches of shallow texture feature extraction and deep semantic feature extraction to fuse feature expressions of different resolutions to form a first multi-level semantic feature covering local details and global semantics. The target domain image set uses the same network architecture, combined with its unique data distribution characteristics, to generate a second multi-level semantic feature that adapts to the spatial structure of the target domain. This heterogeneous design effectively solves the problem of a single feature extraction network being sensitive to domain differences in cross-domain scenarios, and provides a hierarchical feature foundation for subsequent transfer learning.

[0062] The generation of multi-level semantic features relies on a cross-resolution feature fusion strategy. By aligning the dimensions of shallow texture features and channel-wise superposition of deep semantic features, feature representations covering different levels of abstraction are formed. This multi-level feature not only preserves the spatial structural information of source and target domain images but also enhances the feature's adaptability to inter-domain distribution differences. This provides more discriminative input for dynamic domain alignment and spatial transformation, thereby supporting the modeling of complex nonlinear relationships in cross-domain transfer learning and improving the model's robustness to inter-domain distribution differences.

[0063] S130, generating a cross-domain migration weight matrix through a dynamic domain similarity measurement module based on the first multi-level semantic features and the second multi-level semantic features;

[0064] Exemplarily, the dynamic domain similarity measurement module generates a cross-domain migration weight matrix by quantifying the distribution difference of multi-level semantic features between the source domain and the target domain, providing adaptive guidance for feature migration. The module first analyzes the hierarchical distribution characteristics of the first multi-level semantic features (source domain) and the second multi-level semantic features (target domain), and dynamically captures the similarities and differences of features between domains through global correlation measurement and local adaptability evaluation. Based on this, the module constructs cross-level attention weights to form a weight matrix that reflects the importance of migration of different feature levels, thereby achieving enhanced migration of key features and suppression of non-key features, and improving the accuracy of cross-domain feature alignment.

[0065] The dynamic domain similarity measurement module integrates the global distribution trends and local structural details of the source and target domains through an interactive feature matching strategy. Global correlation measurement focuses on the similarity of overall feature distributions between domains, while local adaptability assessment performs fine-grained analysis of differences within specific spatial regions or semantic levels. The synergy between the two enables the generated weight matrix to both adapt to macroscopic distribution differences between domains and finely adjust the strength of local feature migration. This provides an interpretable migration basis for subsequent spatial transformation and channel reorganization, enhancing the model's adaptability to complex domain differences.

[0066] S140, performing spatial transformation on the first multi-level semantic features based on the deformable feature pyramid to generate a migration feature map;

[0067] For example, the deformable feature pyramid achieves nonlinear spatial transformation of multi-level semantic features in the source domain by adaptively adjusting the spatial deformation parameters of the convolution kernel. Guided by the local gradient features of the target domain image, this module dynamically corrects the geometric structure of the convolution kernel to align the spatial distribution of source domain features with that of the target domain. By adjusting the spatial position of the feature map layer by layer, a migration feature map adapted to the target domain scene is generated, effectively alleviating the spatial distribution differences between domains caused by perspective, deformation, or occlusion, and providing spatially consistent feature input for subsequent channel reconstruction.

[0068] The deformable feature pyramid achieves feature alignment at different levels of abstraction through a multi-level structure. Shallow features focus on spatial correction of local details, while deep features focus on adapting global semantic structure. Through cross-level feature fusion, the transferred feature map not only retains the semantic information of the source domain but also incorporates the spatial distribution characteristics of the target domain, thereby enhancing the model's adaptability to complex cross-domain scenarios and providing a robust feature representation foundation for the final classification task.

[0069] S150, based on the cross-domain migration weight matrix, reorganize the channels of the migration feature map to construct the target domain adaptation feature representation;

[0070] For example, the core of channel reorganization is to use the cross-domain migration weight matrix to dynamically adjust the channel importance of the migration feature map to adapt to the data characteristics of the target domain. The cross-domain migration weight matrix reflects the migration priority of different channel features between the source domain and the target domain. The channel attention mechanism is used to weight each channel of the migration feature map, highlighting the feature channels that are highly relevant to the target domain classification task and suppressing redundant or interfering channels. This dynamic channel screening strategy can effectively optimize the semantic adaptability of feature representation, so that the target domain adaptive feature representation focuses more on the key information that can be transferred between domains, providing more discriminative input for subsequent classification tasks.

[0071] The construction of target domain-adaptive feature representations further combines cross-level feature fusion with dimensionality reduction strategies. By splicing and integrating feature information from different levels of abstraction across multiple channels and retaining core semantic features during dimensionality reduction, we ensure that the dimensionality of the final features matches the requirements of the target domain classifier. This process not only achieves refined reorganization of feature channels but also improves the generalization of feature representations in target domain scenarios through cross-domain weight-guided feature optimization, laying the foundation for efficient classifier decision-making.

[0072] S160 : Based on the target domain adaptation feature representation, output the classification result of the target domain image through the target domain classifier.

[0073] For example, the target domain classifier completes the target classification decision in the cross-domain scenario by parsing the semantic information represented by the target domain adaptation feature. The target domain adaptation feature representation integrates the semantic information migrated from the source domain with the spatial adaptation characteristics of the target domain. The classifier performs dimensionality reduction mapping on the features through a fully connected layer and converts them into low-dimensional classification feature vectors. Subsequently, the feature vector is nonlinearly transformed through a preset activation function to generate a probability distribution reflecting the confidence of each category, thereby quantifying the possibility that the target domain image belongs to different categories. This process converts complex cross-domain feature representations into interpretable classification probabilities, providing data support for the final decision.

[0074] The classification result is determined based on the maximum probability distribution principle, which selects the class label with the highest probability as the classification result for the target domain image. This method, through the use of probability thresholds, effectively avoids classification ambiguity caused by differences in feature distribution in cross-domain scenarios, ensuring the reliability and robustness of the output. The final classification output not only demonstrates the discriminative power of the adapted feature representation for the target domain but also verifies the overall effectiveness of the cross-domain transfer learning process, providing high-precision image classification tasks for practical application scenarios.

[0075] In summary, the embodiment of the present application effectively extracts multi-level semantic features of the source domain and the target domain through a heterogeneous feature extraction network, combines the dynamic domain similarity measurement module to adaptively generate a cross-domain migration weight matrix, and uses a deformable feature pyramid to realize feature space transformation and channel reorganization, and finally constructs a target domain adaptation feature representation. This method can accurately quantify the hierarchical differences in cross-domain feature distribution, adaptively adjust the migration weights, significantly improve the fine-grained matching capability of feature alignment, and overcome the limitations of global feature alignment in traditional methods. Through nonlinear spatial transformation and channel attention mechanism, the model can flexibly adapt to the local structural characteristics of the target domain and enhance the robustness to complex domain differences. Compared with the existing technology, while reducing the consumption of computing resources, it improves the classification accuracy in cross-domain scenarios, especially when small samples or domain differences are significant, it shows stronger generalization ability, providing an efficient and robust image classification solution for practical application scenarios.

[0076] In some instances, based on a heterogeneous feature extraction network, first multi-level semantic features of a source domain image set are extracted, including:

[0077] Based on the source domain image set, shallow texture features are extracted by dilating convolutional layers with a preset number in the first convolutional branch of the heterogeneous feature extraction network, wherein the dilation rate of the convolution kernel of the first convolutional branch is negatively correlated with the image resolution;

[0078] Based on the source domain image set, deep semantic features are extracted through the cascaded atrous spatial pyramid pooling module in the second convolutional branch of the heterogeneous feature extraction network, where the cascaded atrous spatial pyramid pooling module includes multiple sets of parallel convolutional layers with different atrous rates;

[0079] Based on shallow texture features and deep semantic features, a first multi-level semantic feature is generated through a cross-resolution stitching operation, wherein the cross-resolution stitching operation includes feature map size alignment and channel dimension superposition.

[0080] Exemplarily, based on the source domain image set, the first convolution branch of the heterogeneous feature extraction network uses a preset number of dilated convolution layers to extract shallow texture features. The dilation rate of the convolution kernel of the dilated convolution layer is negatively correlated with the resolution of the input image, that is, a smaller dilation rate (such as a dilation rate of 2) is used for high-resolution images to increase the effective receptive field of the convolution kernel and capture a larger range of local texture details; a larger dilation rate (such as a dilation rate of 4 or 6) is used for low-resolution images to reduce computational redundancy through sparse sampling while maintaining sensitivity to high-frequency texture features. This design achieves adaptive extraction of texture features at different resolutions by dynamically adjusting the dilation rate, ensuring that shallow features can effectively represent image edges, corners and local texture patterns, and provide data support for subsequent multi-level semantic feature fusion.

[0081] Based on the source domain image set, deep semantic features are extracted through the Cascaded Atrous Spatial Pyramid Pooling (CASPP) module in the second convolutional branch of the heterogeneous feature extraction network. The CASPP module consists of multiple sets of parallel convolutional layers, each with a different dilation rate (e.g., 1, 2, or 4), capturing semantic information of varying granularity within the image using a multi-scale receptive field. Through a cascaded structure, the CASPP module sequentially fuses the outputs of convolutional layers with different dilation rates to form a deep semantic feature representation that covers both global context and local details. For example, convolutional layers with smaller dilation rates focus on local semantic associations, while convolutional layers with larger dilation rates extract long-range dependencies. This multi-scale feature fusion mechanism effectively enhances the semantic expressive power of deep features for complex scenes and provides an abstract hierarchical feature representation for cross-domain transfer.

[0082] Based on the shallow texture features and deep semantic features, the first multi-level semantic features are generated through cross-resolution splicing operations. The cross-resolution splicing operation includes feature map size alignment and channel dimension superposition; wherein, feature map size alignment is to adjust the size of the shallow feature map to the same size as the deep feature map through bilinear interpolation or deconvolution operations. Figure 1The channel dimension superposition is to splice the adjusted shallow texture features and deep semantic features along the channel dimension to form a multi-level feature representation that integrates local details and global semantics. For example, if the size of the shallow feature map is H×W×C1 and the size of the deep feature map is H×W×C2, the size of the feature map generated after splicing is H×W×(C1+C2). This operation not only retains the diversity of features at different levels, but also enhances the multi-level representation ability of features for source domain images through cross-resolution fusion. The first multi-level semantic feature contains both the fine-grained texture information of the source domain image and the high-level semantic abstraction, providing highly discriminative input features for subsequent cross-domain migration weight calculation and spatial transformation.

[0083] In some instances, based on the target domain image set, generating a second multi-level semantic feature having the same dimension as the first multi-level semantic feature through a synchronous processing flow of the first convolution branch and the second convolution branch includes:

[0084] Exemplarily, based on the target domain image set, the generation of the second multi-level semantic features is achieved through the synchronous processing flow of the first convolution branch and the second convolution branch of the heterogeneous feature extraction network. In the first convolution branch, a preset number of dilated convolution layers dynamically adjust the convolution kernel dilation rate according to the resolution of the target domain image. In the second convolution branch, the CASPP module extracts deep semantic features of the target domain image through multiple sets of parallel convolution layers. The CASPP module contains multiple sets of convolution layers with different dilation rates, each of which independently extracts multi-scale semantic information of the target domain image: convolution layers with smaller dilation rates focus on local semantic associations and capture fine-grained target structures; convolution layers with larger dilation rates expand the receptive field and extract global contextual information of the target domain image. Through a cascade structure, the output features of convolution layers with different dilation rates are sequentially fused across scales to form a deep feature representation that covers local details and global semantics. Subsequently, the shallow texture features and deep semantic features are fused through a cross-resolution splicing operation, so that the second multi-level semantic features have both the local texture details and high-level semantic abstraction of the target domain image. Its dimensional structure is consistent with the first multi-level semantic features generated in the source domain, providing a unified feature input basis for subsequent cross-domain migration weight calculation and spatial alignment.

[0085] In some examples, generating a cross-domain transfer weight matrix based on the first multi-level semantic features and the second multi-level semantic features by a dynamic domain similarity measurement module includes:

[0086] Based on the first multi-level semantic features, the cross-level distribution difference measurement is performed through the maximum mean difference calculation module to generate the source domain feature distribution vector;

[0087] Based on the second multi-level semantic features, probability density modeling is performed through the kernel density estimation module to generate the target domain feature distribution vector;

[0088] Based on the source domain feature distribution vector and the target domain feature distribution vector, interactive similarity matching is performed through a dual-channel attention mechanism to generate global attention weights and local attention weights. The global attention weight is generated by calculating the global correlation between the source domain feature distribution vector and the target domain feature distribution vector based on cosine similarity through the first channel of the dual-channel attention mechanism; the local attention weight is generated by calculating the local adaptability of the source domain feature distribution vector and the target domain feature distribution vector based on element-by-element difference through the second channel of the dual-channel attention mechanism;

[0089] Based on the weighted fusion results of global attention weights and local attention weights, a cross-domain transfer weight matrix is generated.

[0090] Exemplarily, the first multi-level semantic feature (source domain) is measured for cross-level distribution difference through the Maximum Mean Discrepancy (MMD) calculation module. The MMD calculation module calculates the distribution difference of semantic features at different levels of the source domain (such as shallow texture features and deep semantic features) in the reproducing kernel Hilbert space (RKHS). Specifically, the feature map of each level is mapped to the sample mean, and the distribution distance between different levels is calculated by a preset kernel function (such as a Gaussian kernel function) to quantify the multi-level distribution characteristics of the source domain features. The distribution difference measurement result is encoded as a source domain feature distribution vector. The source domain feature distribution vector includes the distribution distance weights of the features at each level, reflecting the distribution heterogeneity of features at different abstract levels within the source domain, and providing a basis for subsequent cross-domain similarity matching.

[0091] Based on the second multi-level semantic features (target domain features), non-parametric probability density modeling is first performed through the Kernel Density Estimation (KDE) module. The KDE module performs non-parametric probability density estimation on the feature maps of each level of the target domain, and uses a preset kernel function (such as the Epanechnikov kernel function) to calculate the density distribution of each sample point in the feature space. Specifically, for each feature level, the KDE module calculates the local density distribution of feature points through a sliding window based on the preset kernel bandwidth parameters, and aggregates the density distribution results of each level to generate a target domain feature distribution vector, which characterizes the distribution density of the target domain features in space and the correlation between levels. The target domain feature distribution vector not only reflects the global distribution trend of the target domain features, but also identifies key semantic areas through density peaks, providing a quantitative basis for the target domain feature distribution for cross-domain interactive matching.

[0092] Based on the source domain feature distribution vector and the target domain feature distribution vector, interactive similarity matching is performed through a dual-channel attention mechanism, where the dual-channel attention mechanism includes a first channel (global attention weight generation channel) and a second channel (local attention weight generation channel).

[0093] Global attention weight generation channel: The first channel calculates the global correlation between the source domain feature distribution vector and the target domain feature distribution vector based on the cosine similarity algorithm. Specifically, the source domain feature distribution vector and the target domain feature distribution vector are normalized, and the ratio of their dot product and the modulus length product is calculated to obtain the global attention weight. This weight reflects the degree of match between the source and target domains in terms of overall distribution, guiding the transfer strength of global features in cross-domain transfer.

[0094] Local attention weight generation channel: The second channel calculates the local adaptability of the source domain feature distribution vector and the target domain feature distribution vector based on element-by-element differences. The absolute difference between each element of the source domain feature distribution vector and the target domain feature distribution vector is calculated and mapped to the interval [0, 1] using a preset activation function (such as the Sigmoid function) to generate a local attention weight. This weight identifies the degree of distribution difference between the source and target domains at a specific level or spatial region, allowing for fine-tuning of the migration priority of local features.

[0095] The global attention weights and local attention weights are respectively normalized by the L2 norm to ensure that the weight values are of the same magnitude; according to the preset fusion coefficients (such as the global attention weight α and the local attention weight β, where α+β=1), the normalized global weights and local weights are linearly weighted to generate fusion weights; the fusion weights are expanded according to the hierarchical dimension to construct a cross-domain migration weight matrix corresponding to the multi-level semantic feature channel dimension. Each element of the cross-domain migration weight matrix corresponds to the migration priority of the source domain and target domain features at a specific level or spatial position. A high weight value indicates that the feature has a strong correlation in cross-domain migration and needs to be retained or enhanced first; a low weight value inhibits the migration of non-critical features. Through dynamic weight allocation, the model can adaptively adjust the granularity of cross-domain feature migration, achieve coordinated optimization of global alignment and local adaptation, and improve the accuracy and robustness of cross-domain migration.

[0096] In some instances, spatially transforming the first multi-level semantic features based on the deformable feature pyramid to generate a migration feature map includes:

[0097] Based on the local gradient features of the target domain image set, the deformation parameters of the deformable convolution kernel at each level in the deformable feature pyramid are determined, wherein the local gradient features are extracted through the shallow feature map of the second multi-level semantic features, including:

[0098] Based on the shallow feature map of the second multi-level semantic feature, a directional convolution operation is performed through a preset gradient operator to generate a horizontal gradient component and a vertical gradient component;

[0099] Based on the vector synthesis results of the transverse gradient component and the longitudinal gradient component, the spatial gradient direction distribution of the local gradient feature is determined;

[0100] Based on the statistical histogram of spatial gradient direction distribution, the dominant gradient direction corresponding to each level is extracted through clustering algorithm;

[0101] Based on the angle between the dominant gradient direction and the preset reference direction, the horizontal deformation offset and vertical deformation offset of the deformable convolution kernel of the corresponding level in the deformable feature pyramid are calculated;

[0102] Determine the deformation parameters of the deformable convolution kernel based on the horizontal deformation offset and the vertical deformation offset;

[0103] Based on the deformation parameters, the spatial offset of the deformable convolution kernel is adjusted to generate a deformable convolution kernel that is adaptive to the spatial distribution of the target domain;

[0104] Based on the deformable convolution kernel, bilinear interpolation is performed on the first multi-level semantic features to generate an intermediate feature map after position correction;

[0105] Based on the multi-level structure of the deformable feature pyramid, cross-level feature fusion is performed on the intermediate feature maps to generate a migration feature map that is aligned with the spatial distribution of the target domain.

[0106] Exemplarily, based on the local gradient features of the target domain image set, the deformation parameters of the deformable convolution kernel at each level in the deformable feature pyramid are determined. Specifically, the shallow feature map of the second multi-level semantic feature is used to perform directional convolution operations through a preset gradient operator (such as a Sobel operator) to extract the horizontal gradient component and the vertical gradient component respectively. The horizontal gradient component calculates the grayscale change of the image along the width direction through the horizontal convolution kernel, and the vertical gradient component calculates the grayscale change of the image along the height direction through the vertical convolution kernel. The horizontal and vertical gradient components are vector-synthesized to generate the spatial gradient direction distribution of the local gradient feature, wherein the gradient direction of each pixel is determined by the inverse tangent value of the gradient component, and the gradient amplitude is calculated by the square root of the sum of the squares of the gradient components, thereby quantifying the local structural changes of the target domain image.

[0107] Based on the statistical histogram of the spatial gradient direction distribution, the dominant gradient direction corresponding to each level is extracted through a clustering algorithm (such as K-means or mean shift clustering). Specifically, the gradient direction distribution of the shallow feature map of the target domain is statistically analyzed by histogram, and the cumulative sum of the gradient amplitude in each angle interval is statistically calculated to generate a gradient direction histogram. The peak area in the histogram is divided by a clustering algorithm to extract the dominant gradient direction (such as 0°, 45°, 90°, etc.) of the feature map of each level. This direction reflects the main edge or texture direction of the target domain image at the corresponding level. Based on the angle between the dominant gradient direction and the preset reference direction (such as the horizontal reference direction 0° or the vertical reference direction 90°), the horizontal deformation offset of the deformable convolution kernel of the corresponding level in the deformable feature pyramid is calculated. and vertical deformation offset , generating deformation parameters The deformation parameter reflects the distribution characteristics of the local spatial structure of the target domain and is used to guide the geometric deformation adjustment of the convolution kernel. For example, if the dominant gradient direction is , then the horizontal offset , vertical offset ,in, It is a preset proportional coefficient that controls the correlation between the deformation amplitude and the gradient direction.

[0108] Based on the above deformation parameters , adjust the spatial offset of the deformable convolution kernel to generate a deformable convolution kernel that is adaptive to the spatial distribution of the target domain. Specifically, for each level in the deformable feature pyramid, the deformation offset is mapped to the sampling coordinate grid of the convolution kernel, and the geometric shape of the convolution kernel is adjusted through the coordinate offset operation. For example, the sampling point coordinates of the standard convolution kernel are , the coordinates of the sampling points after deformation are This deformation process enables the convolution kernel to dynamically adapt to the spatial distribution characteristics of the target domain image. For example, the sampling path of the convolution kernel is adjusted according to the edge direction of the target domain, thereby enhancing the adaptability to local deformation or perspective changes. The generation of the deformable convolution kernel is achieved through a parameterized offset layer, whose parameters are dynamically driven by the gradient features of the target domain, ensuring the targeted and adaptive nature of the spatial transformation. The deformed convolution kernel can effectively capture the spatial structural characteristics of the target domain image and provide a geometric adaptation basis for the spatial transformation of the source domain features.

[0109] Based on the deformable convolution kernel, bilinear interpolation is performed on the first multi-level semantic features to generate a position-corrected intermediate feature map. Specifically, for each feature point in the first multi-level semantic features of the source domain, bilinear interpolation sampling is performed in the target domain spatial coordinate system according to the offset of the deformable convolution kernel. The weighted feature value of the offset position is calculated to generate an intermediate feature map that is aligned with the target domain spatial distribution.

[0110] Based on the multi-level structure of the deformable feature pyramid, cross-level feature fusion is performed on the intermediate feature maps to generate a migration feature map that is aligned with the spatial distribution of the target domain. Specifically, the deformable feature pyramid contains multiple levels, and the intermediate feature maps of each level are adjusted to a uniform resolution through upsampling or downsampling operations, and then spliced along the channel dimension. For example, the high-level feature map is superimposed with the low-level feature map through bilinear upsampling to form a composite feature map that integrates multi-scale information. Furthermore, the spliced feature map is subjected to channel compression and nonlinear activation (such as ReLU function) through a 1×1 convolution layer, and finally a migration feature map is generated that is completely aligned with the spatial distribution of the target domain. This feature map not only retains the multi-level semantic information of the source domain, but also achieves high-precision alignment with the spatial distribution of the target domain through spatial correction and cross-level fusion of the deformable convolution kernel, providing robust feature input for subsequent channel reconstruction and classification tasks.

[0111] In summary, through the above steps, the deformable feature pyramid can dynamically adjust the geometric deformation parameters of the convolution kernel, and realize the spatial correction of the source domain features in combination with the local gradient features of the target domain. The generation process of the deformable convolution kernel is based on the target domain gradient direction statistics and cluster analysis, which ensures the physical interpretability of the deformation parameters; bilinear interpolation and cross-level feature fusion further enhance the accuracy and robustness of feature alignment. Compared with the traditional fixed-structure convolution operation, the embodiment of the present application improves the model's adaptability to complex inter-domain spatial differences (such as perspective changes, deformation occlusion), and provides a spatial alignment solution for cross-domain transfer learning.

[0112] In some instances, based on the cross-domain transfer weight matrix, the transferred feature map is channel-reorganized to construct the target domain adaptation feature representation, including:

[0113] Based on the channel dimension of the migration feature map, the channel attention weight is calculated through the cross-domain migration weight matrix to generate the channel attention vector;

[0114] Based on the channel attention vector, each channel feature of the migration feature map is weighted channel by channel to generate a weighted feature map;

[0115] Based on the weighted feature map, cross-level channel splicing is performed through a preset feature fusion strategy to generate a spliced feature map;

[0116] Based on the channel dimension of the spliced feature map, the channel dimension is compressed by a feature compression module with a preset dimensionality reduction ratio to generate a target domain adapted feature representation, where the channel dimension of the target domain adapted feature representation matches the input dimension of the target domain classifier.

[0117] Exemplarily, based on the channel dimension of the migration feature map, the channel attention weight is calculated through the cross-domain migration weight matrix to generate a channel attention vector. Specifically, the dimension of the cross-domain migration weight matrix is C×C', where C is the number of source domain feature channels, C' is the number of target domain feature channels, and each element of the matrix represents the migration priority between the source domain feature channel and the target domain feature channel. The migration feature map (with dimensions of H×W×C, H is the feature map height, W is the feature map width, and C is the number of channels) is expanded into a vector form along the channel dimension, and matrix multiplication is performed with the cross-domain migration weight matrix to generate a channel attention vector (with dimensions of 1×C'). Each element in the channel attention vector corresponds to the importance weight of a specific channel in the target domain classification task. A high weight value indicates that the channel feature has a strong correlation in cross-domain migration.

[0118] Based on the channel attention vector, each channel feature of the migration feature map is weighted channel by channel to generate a weighted feature map. Specifically, each weight value in the channel attention vector is multiplied element by element with its corresponding migration feature map channel to enhance the key channel and suppress the non-key channel. For example, the migration feature map The weight of each channel is , then the weighted feature map The channels are ,in, The original feature map This operation can enhance the contribution of feature channels that are highly relevant to the target domain classification task, while suppressing information from redundant or interfering channels. For example, if the attention weight of a channel is close to 0, its corresponding feature map will be weakened in subsequent processing; conversely, the feature information of high-weight channels will be significantly enhanced. Through channel-by-channel weighting, the semantic expression of the transferred feature map is more focused on the key information for target domain adaptation, laying the foundation for feature fusion.

[0119] Based on weighted feature maps, a pre-set feature fusion strategy is used to perform cross-level channel splicing to generate a spliced feature map. This pre-set feature fusion strategy involves resizing weighted feature maps from different layers (e.g., shallow weighted feature maps of size H1×W1×C1 and deep weighted feature maps of size H2×W2×C2) to a uniform spatial resolution (H×W) through bilinear interpolation or deconvolution. The aligned feature maps are then spliced along the channel dimension to form a composite feature map of dimensions H×W×(C1+C2). This operation, through cross-level feature fusion, integrates feature information from different levels of abstraction, preserving the complementarity between local details and global semantics. Furthermore, the spliced feature maps undergo preliminary channel interaction (e.g., dimensionality reduction or nonlinear activation) through a 1×1 convolutional layer to enhance the expressive power of the features. This cross-level splicing strategy ensures the multi-scale nature of the target domain-adapted feature representation and improves the robustness of the features to complex domain differences.

[0120] Based on the channel dimension of the spliced feature map, the channel dimension is compressed by a feature compression module with a preset dimensionality reduction ratio to generate a target domain adaptation feature representation. The feature compression module consists of a fully connected layer or a 1×1 convolutional layer, and its dimensionality reduction ratio is preset according to the input dimension of the target domain classifier. For example, if the number of channels of the spliced feature map is 1024 and the input dimension of the target domain classifier is 512, the dimensionality reduction ratio is 0.5, and the number of channels is compressed to 512 through the fully connected layer. During the compression process, L2 regularization is used to constrain the weight parameters to avoid overfitting. The channel dimension of the target domain adaptation feature representation finally generated completely matches the classifier input, and its dimension is H×W×D (D is the number of input channels of the target domain classifier). It not only retains the semantic information after cross-level fusion, but also reduces redundancy through dimensionality reduction, ensuring efficient calculation and decision accuracy of the classifier.

[0121] Through the above steps, the channel reorganization process achieves dynamic application of cross-domain transfer weights and optimized adaptation of feature dimensions. The channel attention mechanism strengthens key features and suppresses redundant information through weight allocation; the cross-level splicing strategy integrates multi-granularity semantics to enhance the robustness of feature representation; and the feature compression module balances computational efficiency and classification accuracy through dimensionality reduction. The resulting target domain-adapted feature representation is both discriminative and adaptable, improving the model's classification performance in target domain scenarios.

[0122] In some instances, based on the target domain adapted feature representation, a target domain classifier is used to output a classification result of the target domain image, including:

[0123] Based on the channel dimension of the target domain adaptation feature representation, the feature dimension reduction mapping is performed through the fully connected layer in the target domain classifier to generate a classification feature vector;

[0124] Based on the classification feature vector, a nonlinear transformation is performed through a preset activation function to generate a category probability distribution;

[0125] The classification result of the target domain image is determined based on the category label corresponding to the maximum probability value in the category probability distribution.

[0126] Exemplarily, based on the channel dimension of the target domain adaptation feature representation, the fully connected layer in the target domain classifier is used to perform feature dimensionality reduction mapping to generate a classification feature vector. The input dimension of the fully connected layer strictly matches the channel dimension of the target domain adaptation feature representation, and its output dimension is determined according to the preset number of classification categories. For example, the dimension of the target domain adaptation feature representation is H×W×D, and the fully connected layer linearly transforms the input features along the channel dimension through the weight matrix and maps them to a dimension that matches the number of target categories. For example, if the number of target categories is K, the dimension of the weight matrix of the fully connected layer is D×K, and the input features are converted into a 1×1×D vector after global average pooling, and multiplied by the weight matrix to generate a 1×K classification feature vector. This process extracts key semantic information in the target domain adaptation feature representation through linear combination, providing a compact feature expression for subsequent probability distribution generation.

[0127] Based on the classification feature vector, a nonlinear transformation is performed using a preset activation function to generate a probability distribution reflecting the confidence level of each category. The preset activation function is the Softmax function, which takes each element of the classification feature vector as input and outputs a normalized probability value. The Softmax function uses exponential calculations and normalization to convert the raw scores of the classification feature vector into a probability distribution in the interval [0, 1], ensuring that the sum of all class probabilities is 1. This probability distribution quantifies the likelihood that the target domain image belongs to each category, providing an interpretable numerical basis for classification decisions.

[0128] Based on the category label corresponding to the maximum probability value in the category probability distribution, the classification result of the target domain image is determined. Specifically, the probability distribution output by the Softmax function is retrieved for the maximum value, and the category index with the highest probability value is selected as the final classification label. For example, if the probability distribution is [0.05, 0.75, 0.20], the category with index 1 (category index starts from 0) is selected as the classification result. This application effectively avoids the classification ambiguity problem caused by feature distribution differences or noise interference in cross-domain scenarios through a probability threshold judgment mechanism, ensuring the uniqueness and reliability of the output results. The determination of the final classification label not only depends on the discriminative ability of the adapted feature representation of the target domain, but also verifies the effectiveness of the cross-domain transfer learning process through quantitative analysis of the probability distribution, providing high-precision classification output for practical applications.

[0129] See also Figure 2 , which is a schematic diagram of the structure of an image classification device for cross-domain transfer learning provided in an embodiment of the present application, comprising:

[0130] An image data acquisition unit 21 is configured to acquire a source domain image set and a target domain image set;

[0131] A semantic feature extraction unit 22 extracts first multi-level semantic features of the source domain image set and second multi-level semantic features of the target domain image set based on a heterogeneous feature extraction network;

[0132] A weight matrix generating unit 23 generates a cross-domain migration weight matrix based on the first multi-level semantic features and the second multi-level semantic features through a dynamic domain similarity measurement module;

[0133] A feature map generating unit 24 performs spatial transformation on the first multi-level semantic features based on the deformable feature pyramid to generate a migration feature map;

[0134] The feature representation construction unit 25 performs channel reorganization on the migration feature map based on the cross-domain migration weight matrix to construct a target domain adaptation feature representation;

[0135] The image classification output unit 26 outputs the classification result of the target domain image through the target domain classifier based on the target domain adaptation feature representation.

[0136] See also Figure 3 An embodiment of the present application also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored on the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of any method of image classification of cross-domain transfer learning.

[0137] Since the electronic device introduced in this embodiment is a device used to implement an image classification device for cross-domain transfer learning in the embodiment of this application, based on the method introduced in the embodiment of this application, technical personnel in this field can understand the specific implementation of the electronic device of this embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of this application will not be introduced in detail here. As long as the equipment used by technical personnel in this field to implement the method in the embodiment of this application falls within the scope of protection to be protected by this application.

[0138] During the specific implementation process, when the computer program 311 is executed by the processor, any implementation method of the embodiments corresponding to the first aspect can be implemented.

[0139] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0140] Those skilled in the art will appreciate that embodiments of the present application may provide methods, systems, or computer program products. Thus, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-readable storage media containing computer-readable program code.

[0141] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0142] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0144] The present application also provides a computer program product, which includes computer software instructions. When the computer software instructions are executed on a processing device, the processing device executes Figure 1 A process of an image classification method for cross-domain transfer learning in the corresponding embodiment.

[0145] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium can be a magnetic medium, an optical medium or a semiconductor medium, etc.

[0146] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0147] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0148] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0149] In addition, the functional units in the various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The above-mentioned integrated units may be implemented in the form of hardware and / or software functional units.

[0150] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device to execute all or part of the steps of the various embodiments of the method of the present application.

[0151] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

[0152] Although the preferred embodiments of this specification have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of this specification.

[0153] Obviously, those skilled in the art may make various changes and modifications to this specification without departing from the spirit and scope of this specification. Thus, if such changes and modifications fall within the scope of the claims of this specification and their equivalents, this specification is intended to include such changes and modifications.

Claims

1. A cross-domain transfer learning image classification method, characterized in that: include: Obtain a source domain image set and a target domain image set; Extracting first multi-level semantic features of the source domain image set and second multi-level semantic features of the target domain image set based on a heterogeneous feature extraction network; Based on the first multi-level semantic features and the second multi-level semantic features, generating a cross-domain migration weight matrix through a dynamic domain similarity measurement module, including: Based on the first multi-level semantic features, a cross-level distribution difference measurement is performed through a maximum mean difference calculation module to generate a source domain feature distribution vector; Based on the second multi-level semantic features, probability density modeling is performed through a kernel density estimation module to generate a target domain feature distribution vector; Based on the source domain feature distribution vector and the target domain feature distribution vector, interactive similarity matching is performed through a dual-channel attention mechanism to generate a global attention weight and a local attention weight, wherein the global attention weight is generated by calculating the global correlation between the source domain feature distribution vector and the target domain feature distribution vector based on cosine similarity through the first channel of the dual-channel attention mechanism; the local attention weight is generated by calculating the local adaptability of the source domain feature distribution vector and the target domain feature distribution vector based on element-by-element difference through the second channel of the dual-channel attention mechanism; Generate a cross-domain migration weight matrix based on the weighted fusion result of the global attention weight and the local attention weight; Based on the deformable feature pyramid, spatially transforming the first multi-level semantic features to generate a migration feature map includes: Determining deformation parameters of a deformable convolution kernel at each level in a deformable feature pyramid based on local gradient features of the target domain image set, wherein the local gradient features are extracted through a shallow feature map of the second multi-level semantic features, comprising: Based on the shallow feature map of the second multi-level semantic feature, a directional convolution operation is performed using a preset gradient operator to generate a horizontal gradient component and a vertical gradient component; Determining the spatial gradient direction distribution of the local gradient feature based on a vector synthesis result of the transverse gradient component and the longitudinal gradient component; Based on the statistical histogram of the spatial gradient direction distribution, the dominant gradient direction corresponding to each level is extracted by a clustering algorithm; Calculating a horizontal deformation offset and a vertical deformation offset of the deformable convolution kernel of a corresponding level in the deformable feature pyramid based on an angle between the dominant gradient direction and a preset reference direction; Determining a deformation parameter of the deformable convolution kernel based on the horizontal deformation offset and the vertical deformation offset; Based on the deformation parameter, adjusting the spatial offset of the deformable convolution kernel to generate a deformable convolution kernel that is adaptive to the spatial distribution of the target domain; Based on the deformable convolution kernel, performing bilinear interpolation calculation on the first multi-level semantic features to generate an intermediate feature map after position correction; Based on the multi-level structure of the deformable feature pyramid, cross-level feature fusion is performed on the intermediate feature map to generate a migration feature map aligned with the spatial distribution of the target domain; Based on the cross-domain migration weight matrix, the migration feature map is channel-reorganized to construct a target domain adaptation feature representation; Based on the target domain adaptation feature representation, a classification result of the target domain image is outputted through a target domain classifier.

2. The method according to claim 1, characterized in that The step of extracting first multi-level semantic features of the source domain image set based on a heterogeneous feature extraction network includes: Based on the source domain image set, extracting shallow texture features through a preset number of dilated convolutional layers in a first convolutional branch of a heterogeneous feature extraction network, wherein a convolution kernel dilation rate of the first convolutional branch is negatively correlated with the image resolution; Based on the source domain image set, extracting deep semantic features through a cascaded atrous spatial pyramid pooling module in the second convolutional branch of the heterogeneous feature extraction network, wherein the cascaded atrous spatial pyramid pooling module includes multiple groups of parallel convolutional layers with different atrous rates; Based on the shallow texture features and the deep semantic features, the first multi-level semantic features are generated through a cross-resolution stitching operation, wherein the cross-resolution stitching operation includes feature map size alignment and channel dimension superposition.

3. The method according to claim 1, characterized in that The step of performing channel reorganization on the migration feature map based on the cross-domain migration weight matrix to construct a target domain adaptation feature representation includes: Based on the channel dimension of the migration feature map, channel attention weight is calculated using the cross-domain migration weight matrix to generate a channel attention vector; Based on the channel attention vector, performing a channel-by-channel weighting operation on each channel feature of the migration feature map to generate a weighted feature map; Based on the weighted feature map, cross-level channel splicing is performed through a preset feature fusion strategy to generate a spliced feature map; Based on the channel dimension of the spliced feature map, a feature compression module with a preset dimensionality reduction ratio is used to perform channel dimension compression processing to generate a target domain adapted feature representation, wherein the channel dimension of the target domain adapted feature representation matches the input dimension of the target domain classifier.

4. The method according to claim 1, wherein Outputting a classification result of a target domain image through a target domain classifier based on the target domain adaptation feature representation includes: Based on the channel dimension of the target domain adaptation feature representation, feature dimensionality reduction mapping is performed through a fully connected layer in the target domain classifier to generate a classification feature vector; Based on the classification feature vector, a nonlinear transformation is performed through a preset activation function to generate a category probability distribution; Based on the category label corresponding to the maximum probability value in the category probability distribution, a classification result of the target domain image is determined.

5. An image classification device for cross-domain transfer learning, characterized in that: include: An image data acquisition unit, configured to acquire a source domain image set and a target domain image set; a semantic feature extraction unit, which extracts first multi-level semantic features of the source domain image set and second multi-level semantic features of the target domain image set based on a heterogeneous feature extraction network; A weight matrix generating unit generates a cross-domain migration weight matrix based on the first multi-level semantic features and the second multi-level semantic features through a dynamic domain similarity measurement module, including: Based on the first multi-level semantic features, a cross-level distribution difference measurement is performed through a maximum mean difference calculation module to generate a source domain feature distribution vector; Based on the second multi-level semantic features, probability density modeling is performed through a kernel density estimation module to generate a target domain feature distribution vector; Based on the source domain feature distribution vector and the target domain feature distribution vector, interactive similarity matching is performed through a dual-channel attention mechanism to generate a global attention weight and a local attention weight, wherein the global attention weight is generated by calculating the global correlation between the source domain feature distribution vector and the target domain feature distribution vector based on cosine similarity through the first channel of the dual-channel attention mechanism; the local attention weight is generated by calculating the local adaptability of the source domain feature distribution vector and the target domain feature distribution vector based on element-by-element difference through the second channel of the dual-channel attention mechanism; Generate a cross-domain migration weight matrix based on the weighted fusion result of the global attention weight and the local attention weight; The feature map generation unit performs spatial transformation on the first multi-level semantic features based on the deformable feature pyramid to generate a migration feature map, including: Determining deformation parameters of a deformable convolution kernel at each level in a deformable feature pyramid based on local gradient features of the target domain image set, wherein the local gradient features are extracted through a shallow feature map of the second multi-level semantic features, comprising: Based on the shallow feature map of the second multi-level semantic feature, a directional convolution operation is performed using a preset gradient operator to generate a horizontal gradient component and a vertical gradient component; Determining the spatial gradient direction distribution of the local gradient feature based on a vector synthesis result of the transverse gradient component and the longitudinal gradient component; Based on the statistical histogram of the spatial gradient direction distribution, the dominant gradient direction corresponding to each level is extracted by a clustering algorithm; Calculating a horizontal deformation offset and a vertical deformation offset of the deformable convolution kernel of a corresponding level in the deformable feature pyramid based on an angle between the dominant gradient direction and a preset reference direction; Determining a deformation parameter of the deformable convolution kernel based on the horizontal deformation offset and the vertical deformation offset; Based on the deformation parameter, adjusting the spatial offset of the deformable convolution kernel to generate a deformable convolution kernel that is adaptive to the spatial distribution of the target domain; Based on the deformable convolution kernel, performing bilinear interpolation calculation on the first multi-level semantic features to generate an intermediate feature map after position correction; Based on the multi-level structure of the deformable feature pyramid, cross-level feature fusion is performed on the intermediate feature map to generate a migration feature map aligned with the spatial distribution of the target domain; A feature representation construction unit, which performs channel reorganization on the migration feature map based on the cross-domain migration weight matrix to construct a target domain adaptation feature representation; The image classification output unit outputs the classification result of the target domain image through the target domain classifier based on the target domain adaptation feature representation.

6. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein the processor is configured to implement the steps of the image classification method for cross-domain transfer learning according to any one of claims 1 to 4 when executing the computer program stored in the memory.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image classification method of cross-domain transfer learning according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Method for detecting X-ray mammary gland lesion image based on feature pyramid network under transfer learning

    CN110674866A

  • Method for quickly mixing high-order attention domain adversarial network based on transfer learning

    CN112446423A

  • Progressive information decoupling-based cross-domain model training method

    CN116778277A