Unsupervised domain-adaptive image classification method based on ViT

CN120182710BActive Publication Date: 2026-09-01CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326308.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-09-01
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

[0006]本发明要解决的技术问题是提供一种基于ViT的无监督域适应图像分类方法,以解决现有无监督域适应图像分类方法中缺乏对领域级信息捕获和局部信息交互的问题

Benefits of technology

[0036]本发明的有益效果是,本发明提供了一种基于ViT的无监督域适应图像分类方法,形成了跨领域的长程依赖关系,而不是显式地对域差异进行建模。在ViT框架下引入了基于全局对应关系的域级交叉注意力机制,以学习具有跨域迁移能力的特征。该方法直接在领域级进行操作,有助于建立不同领域之间的样本对应关系。引入卷积-transformer特征交互模块来将局部特征引入视觉Transformer,实现全局信息和局部信息的交互和整合,在不改变ViT结构的前提下引入卷积特征。同时,通过双向交互,在进一步增强视觉里程计语义表征的同时,缓解了视觉transformer中缺少局部信息和多尺度特征交互的问题。弥合了基于patch的ViT处理和卷积处理之间的鸿沟,这种交互方法允许利用两种模型的优点。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182710B_ABST
    Figure CN120182710B_ABST
Patent Text Reader

Abstract

This invention provides an unsupervised domain-adaptive image classification method based on ViT. It introduces a domain-cross-attention module based on global correspondence into the original ViT to interact with sample features from different domains at the domain level, capturing the global correspondence between cross-domain sample features. Global features are obtained based on the captured global correspondence between cross-domain sample features. Local features are obtained by introducing a feature interaction module into the original ViT. Local features interact and are superimposed with global features. The original ViT, after incorporating the domain-cross-attention module and the convolutional-transformer feature interaction module, forms a fused ViT. The fused ViT is trained to obtain a trained fused ViT. The trained fused ViT is then used to perform image classification tasks. This solves the problem of existing unsupervised domain-adaptive image classification methods lacking the capture of domain-level information and the interaction of local information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image classification, specifically relating to an unsupervised domain-adaptive image classification method based on ViT. Background Technology

[0002] Deep learning has made remarkable progress in various recognition tasks, especially in image classification. Unsupervised domain adaptation (UDA) image classification methods primarily involve the fields of computer vision, machine learning, and image processing. It focuses on how a model trained on a source domain can maintain good performance on a target domain with similar features but different distributions, without requiring label information from the target domain. In computer vision (CV), a large number of methods have been extensively studied. These include Convolutional Neural Networks (CNNs) and, more recently, Vision Transformers (ViTs). ViTs have shown significant generalization ability in supervised learning scenarios, often outperforming CNN-based techniques. Despite these advances, Deep Neural Networks (DNNs) address the phenomenon of domain transfer, where models are trained on data with different distributions in the source and target datasets. Traditional solutions involve labeling the target domain data, which is not always feasible due to associated costs. Unsupervised domain adaptation (UDA) has emerged as a strategic choice, relying on a labeled source dataset and an unlabeled target dataset, often exhibiting significant domain differences. UDA aims to bridge the gap between these two domains, thereby improving model performance on the target dataset. Class-level alignment-based solutions reduce the domain gap by generating pseudo-labels for target samples, such as metric learning, adversarial training, and optimal transfer. Furthermore, some works have explored the potential of ViT in non-trivial UDA tasks, achieving results surpassing CNNs.

[0003] Visual Transformer (ViT) is a pioneering work applying convolution-free transform structures to image classification. Several works have already applied ViT to the unsupervised domain adaptation domain. TVT proposes an adaptive module to capture transferable and discriminative features of the domain data. SSRT proposes a framework with a transformer backbone and a secure self-refining strategy to handle the problem in the case of large domain gaps. CDTrans proposes a two-step framework that utilizes cross-attention from ViT for direct feature alignment and uses pre-generated pseudo-labels on the target samples. PMTrans interprets the unsupervised domain adaptation process as a minimax game and finds an intermediate domain.

[0004] From the perspective of domain adaptation, rigorous transfer theory states that the target error of a classifier is primarily bounded by its source error plus the inter-domain divergence. Guided by this theory, a series of CNN-based methods attempt to learn domain-invariant features by explicitly minimizing the source error and the statistical distribution differences between domains. The tremendous success of the Transformer in computer vision has further inspired recent advances in UDA. However, simply applying the Transformer can only establish the interaction relationships between image patches and cannot characterize the relationships between cross-domain samples. Therefore, learning only patch-level information is insufficient.

[0005] Secondly, regarding feature representation, to process 2D images using the Transformer, the input image is first converted into a sequence of tokens corresponding to image patches. Then, the attention module focuses on all the tokens and computes a weighted sum as the tokens for the next layer. This effectively expands the receptive field to the entire image through a single self-attention layer. However, a problem with the visual Transformer conformational transformation lies in the contradiction between global connectivity and the convolutional philosophy. Considering the advantages of CNNs over Transformers, a natural question arises: can the locality of CNNs be effectively combined with the global connectivity of the visual transformer to improve performance, rather than increasing model complexity? Summary of the Invention

[0006] The technical problem to be solved by this invention is to provide an unsupervised domain-adaptive image classification method based on ViT, so as to solve the problem that existing unsupervised domain-adaptive image classification methods lack domain-level information capture and local information interaction.

[0007] To achieve the above objectives, one or more embodiments of this application provide an unsupervised domain-adaptive image classification method based on ViT, which includes the following steps:

[0008] Step 1: Introduce a domain-based cross-attention module based on global correspondence into the original ViT to interact with sample features from different domains at the domain level and capture the global correspondence between cross-domain sample features; obtain sample-level global features based on the captured global correspondence between cross-domain sample features, input the sample-level global features into the original ViT, and the original ViT outputs multi-level global features.

[0009] Step 2: In the original ViT, a convolution-transformer feature interaction module is introduced to input multi-level global features into the convolution-transformer feature interaction module to obtain local features; the obtained local features are first interacted with the multi-level global features, and then superimposed to obtain fused features;

[0010] Step 3: After introducing the domain cross-attention module and the convolution-transformer feature interaction module, the original ViT is transformed into a fused ViT. The multi-level global features obtained in Step 1 are used as the input of the fused ViT, and the fused features obtained in Step 2 are used as the output of the fused ViT. The fused ViT is iteratively trained using adversarial training to obtain the trained fused ViT.

[0011] Step four: Use the trained ViT to perform the image classification task.

[0012] Based on the above technical solution of the present invention, the following improvements can also be made:

[0013] Optionally, in step one, the global correspondence between cross-domain sample features is represented by an attention score matrix. The attention score matrix is ​​calculated based on the query generated from the source domain sample features and the key value generated from the target domain sample features. The target domain features are reconstructed based on the attention score matrix. The reconstructed target domain features continue to use the labels of the source domain features and are used as new source domain features, i.e., sample-level global features. The new source domain features are input into the original ViT, and the original ViT outputs multi-level global features.

[0014] Optionally, in step two, the process of obtaining local features by the convolution-transformer feature interaction module is as follows:

[0015] First, the input image is divided into several image blocks, and the flattened feature sequence of ViT is projected back into the 2D image space using the orthographic projection operator;

[0016] Then, the two-dimensional convolution operator is repeatedly applied to the local blocks in the flattened feature sequence of ViT projected back to 2D space to generate local convolution features.

[0017] For the obtained local convolutional features, the inverse projection operator is used to project them back to the original dimension in the feature space to obtain new features, which are local features.

[0018] Optionally, in the two-dimensional convolution operator, the kernel size is k, and zero padding of size (k-1) / 2 is added around the perimeter of the region covered by the kernel to keep the local convolution feature size unchanged.

[0019] Optionally, the process of obtaining the fused features is as follows: After obtaining the local features, the local features and multi-level global features are first interacted using a self-attention mechanism, and then the local features and multi-level global features are superimposed using residual connections to obtain the fused features.

[0020] Optionally, before step three is executed, the merged ViT is divided into several stages, with each stage handling patches of different sizes.

[0021] Optionally, the loss function for adversarial training is as follows:

[0022]

[0023] Among them, the i-th sample and its corresponding one-hot tag Represents a labeled source set Using the j-th sample Represents the unlabeled target set n s and n t Let D represent the sample sizes in the source and target domains, respectively. s and D t The union of d i Let M be the domain label, λ be the model's prediction result for the input sample, F be the weighting coefficient, and C be the classifier. It is the cross-entropy loss on labeled source domain samples. It is the classic domain alignment loss based on domain adversarial learning:

[0024]

[0025] in, Represents the i-th sample in the source domain. Let x represent the i-th sample in the target domain. i Let y represent the i-th input sample. i This represents the true label of the i-th sample in the source domain.

[0026] According to a second aspect of the present invention, an unsupervised domain-adaptive image classification system based on ViT is provided, the system comprising a ViT architecture and:

[0027] The domain cross-attention module introduces a global correspondence-based domain cross-attention module into the original ViT. It interacts with sample features from different domains at the domain level, captures the global correspondence between cross-domain sample features, and obtains global features.

[0028] The convolution-transformer feature interaction module is introduced into the original ViT to obtain local features; the obtained local features are interacted with and superimposed with global features to obtain fused features;

[0029] The training module integrates the domain cross-attention module and the convolution-transformer feature interaction module into ViT to form a fused ViT. The input of the fused ViT is global features, and the output is fused features. The fused ViT is iteratively trained using an adversarial training method to obtain a trained fused ViT.

[0030] The execution module performs image classification tasks based on the trained fusion ViT.

[0031] According to a third aspect of the present invention, an electronic device is provided, the device comprising:

[0032] Memory containing executable program code;

[0033] A processor coupled to the memory;

[0034] The processor calls the executable program code stored in the memory to execute the steps of the ViT-based unsupervised domain adaptation image classification method described above.

[0035] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, which, when invoked, are used to perform the steps of the above-described ViT-based unsupervised domain adaptation image classification method.

[0036] The beneficial effects of this invention are that it provides an unsupervised domain-adaptive image classification method based on ViT, forming long-range dependencies across domains instead of explicitly modeling domain differences. A domain-level cross-attention mechanism based on global correspondence is introduced within the ViT framework to learn features with cross-domain transfer capabilities. This method operates directly at the domain level, facilitating the establishment of sample correspondences between different domains. A convolution-transformer feature interaction module is introduced to introduce local features into the visual Transformer, achieving interaction and integration of global and local information, and introducing convolutional features without altering the ViT structure. Simultaneously, through bidirectional interaction, the semantic representation of visual odometry is further enhanced while alleviating the lack of local information and multi-scale feature interaction in the visual transformer. This method bridges the gap between patch-based ViT processing and convolutional processing, allowing the utilization of the advantages of both models. Attached Figure Description

[0037] Figure 1 This is a flowchart illustrating an unsupervised domain-adaptive image classification method based on ViT, according to an embodiment of the present invention.

[0038] Figure 2This is a schematic diagram of the attention process in an unsupervised domain-adaptive image classification method based on ViT, according to an embodiment of the present invention.

[0039] Figure 3 This is a schematic diagram of feature transformation for an unsupervised domain-adaptive image classification method based on ViT, according to an embodiment of the present invention. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0041] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in one or more embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0042] refer to Figures 1-3 This application discloses one or more embodiments of an unsupervised domain-adaptive image classification method based on ViT, which includes:

[0043] like Figure 1 As shown, this embodiment actually provides an unsupervised domain adaptation image classification method based on ViT, which fuses domain attention and local features. In UDA (unsupervised domain adaptation), the i-th sample is used... and its corresponding one-hot tag Represents a labeled source set Using the j-th sample Represents the unlabeled target set n s and n t Let represent the sample sizes in the source and target domains, respectively. Note that the data in the two domains are sampled from two different distributions, and assume that the two domains share the same label space. The goal is to resolve significant domain divergence and smoothly transfer knowledge from the source domain to the target domain.

[0044] To extend the Transformer to image-based learning tasks, ViT (Vision Transformer) and its variants typically segment each image into N image patches and represent them as feature moments. Applying different projections to Z yields Q, K, and d and d n Indicate their dimensions. The self-attention mechanism in ViT is expressed as:

[0045]

[0046] Attention(Z) = AV

[0047] This formula establishes long-range dependencies between each image patch, enabling ViT to identify image regions relevant to classification. Furthermore, the formula introduces a cross-attention mechanism when Q and K are computed from different images.

[0048] First, we propose a Domain-Level Cross-Attention Module (DoM), which forms long-range dependencies across domains instead of explicitly modeling domain differences. Specifically, we introduce a Domain-Level Cross-Attention Module based on global correspondences into the original ViT framework to learn features with cross-domain transferability. This method operates directly at the domain level, which helps establish sample correspondences between different domains. Unlike traditional ViT-based UDA methods that interact at the patch level or explicitly model distribution differences, we introduce a Domain Attention Module to interact with features from different domains at the domain level, directly characterizing cross-domain sample correspondences and capturing global correspondences between cross-domain sample features, thereby learning transferable features. These global features are obtained by the Transformer Encoder. The purpose of the Domain-Level Cross-Attention Module is to leverage the advantages of long-range correspondences in the attention mechanism to characterize correspondences between cross-domain samples.

[0049] Under the UDA assumption, samples from different domains may have potentially significant distributional differences due to domain shifts. In this embodiment, we refer to the similarity of samples in the same latent class across different domains as global correspondence. Furthermore, we construct domain-level sample correspondences by proposing a domain attention module, where queries originate from the source domain and key-value pairs originate from the target domain. To simulate independent and learnable feature embeddings for queries and key-value pairs, we use projection operators for projection. The projection operators for the source and target domains are denoted as Proj, respectively. s (·) and Proj t (·). Queries and key (value) representations can be:

[0050]

[0051] in, and It is the output of the i-th stage of ViT.

[0052] like Figure 2 As shown, DoM uses a domain-level attention mechanism to exchange sample features across different domains to form an at attention graph. (Also known as the attention score matrix), serving as cross-domain expression coefficients:

[0053]

[0054] The scaled similarity matrix A represents the source sample X s and target sample X t Pairwise similarity in feature space. Reconstructing target domain features F based on attention score matrix A. t A higher attention score indicates that the target domain features are more similar to the source domain features, and the target domain features will receive more attention. (Reconstructed target domain features) It can be represented as:

[0055]

[0056] The reconstructed features are a weighted sum of keys, and the more similar the target domain features are to the source domain features, the higher their attention will be. Although there are differences between the source and target domains, features from different domains of the same category are expected to be more similar, reflecting global correspondences and correspondences between samples of the same category. Because It is based on similarity reconstruction, therefore You can continue to use F s The tag, and As new source domain features, the module obtains new source domain features, i.e., sample-level global features, through the cross-attention mechanism between source and target domain features. The new source domain features will replace the original source domain features in the original ViT and subsequent stages. Specifically, the new source domain features will replace the source domain features of the previous stage and enter the original ViT to obtain multi-level global features.

[0057] Secondly, we introduce a Convolutional-Transformer Feature Interaction Module (CoT) to incorporate local features into ViT, enabling the interaction and integration of global and local information. Specifically, local features are obtained through 2D convolution operators in the interaction module. These local features are then superimposed on the multi-level global features output by the original ViT. Because the local and multi-level global features originate from different architectures, a self-attention mechanism is employed for integration and interaction, serving as input for the next stage.

[0058] As shown in Figure 3, it introduces convolutional features without altering the ViT structure. Simultaneously, through bidirectional interaction, it further enhances the semantic representation of visual odometry while alleviating the problems of lacking local information and multi-scale feature interaction in ViT.

[0059] In a visual converter, an input image of size H×W is divided into image patches of size P×P, and the number of image patches is N = HW / P. 2 To obtain convolutional features, we first use the orthographic projection operator to project the flattened feature sequence of ViT back into the 2D image space.

[0060]

[0061] Where D is the feature dimension, C is the number of color channels in the image, and B represents the number of data samples input into the model at one time.

[0062] Then, for X ′ The two-dimensional convolution operator is repeatedly applied to local patches to generate local convolutional features. This is achieved by sliding the convolution kernel across the patch, computing the dot product at each location, and generating a feature map that captures the spatial hierarchy and local patterns within the patch.

[0063]

[0064] The convolution kernel size is k (k≥3), and zero padding of size (k-1) / 2 is added around the perimeter of the small blocks. It is worth noting that the zero padding here is important for maintaining the feature size.

[0065] For the obtained local convolutional features, we further employ an inverse projection operator in the feature space to project the features back to the original dimension:

[0066]

[0067] The newly obtained feature map F, after being processed by the convolution operator, retains the local information of the image to the greatest extent. However, due to differences in architecture, F and Transformer features exhibit differences in modal representation. To address this issue, we employ a self-attention mechanism to unify convolutional and transform features, enhancing the invariance of representation to modal differences. Here, the transform feature is the Transformer feature, which is a multi-level global feature, representing the features obtained after processing the new source domain features through the original ViT; the convolutional feature, on the other hand, is a local feature.

[0068] Through CoT, we effectively bridge the gap between patch-based ViT processing and convolutional processing. This interactive approach allows us to leverage the strengths of both models. To further facilitate the interaction between convolutional and Transformer features, we utilize residual connections to fuse the outputs of CoT and ViT layers, addressing the weaknesses of ViT in lacking local information and multi-scale features. By combining global and local information, we further enrich and enhance the semantic information of the output features.

[0069] After integrating the DoM and CoT modules into ViT, as shown in Figure 1, ViT is divided into N stages, each with a different patch size.

[0070] For the training process, we use a relatively traditional adversarial training method, with the following loss function:

[0071]

[0072] Among them, the i-th sample and its corresponding one-hot tag Represents a labeled source set Using the j-th sample Represents the unlabeled target set n s and n t Let D represent the sample sizes in the source and target domains, respectively. s and D t The union of d i Let M be the domain label, λ be the model's prediction result for the input sample, F be the weighting coefficient, and C be the classifier. It is the cross-entropy loss on labeled source domain samples. It is the classic domain alignment loss based on domain adversarial learning:

[0073]

[0074] in, Represents the i-th sample in the source domain. Let x represent the i-th sample in the target domain. i Let y represent the i-th input sample. i This represents the true label of the i-th sample in the source domain.

[0075] In another embodiment, an unsupervised domain-adaptive image classification system based on ViT is provided, the system comprising a ViT architecture and:

[0076] The domain cross-attention module introduces a global correspondence-based domain cross-attention module into the original ViT. It interacts with sample features from different domains at the domain level, captures the global correspondence between cross-domain sample features, and obtains global features.

[0077] The convolution-transformer feature interaction module is introduced into the original ViT to obtain local features; the obtained local features are interacted with and superimposed with global features to obtain fused features;

[0078] The training module integrates the domain cross-attention module and the convolution-transformer feature interaction module into ViT to form a fused ViT. The input of the fused ViT is global features, and the output is fused features. The fused ViT is trained using an adversarial training method to obtain a trained fused ViT.

[0079] The execution module performs image classification tasks based on the trained fusion ViT.

[0080] In another embodiment, an electronic device is provided, the device comprising:

[0081] Memory containing executable program code;

[0082] A processor coupled to the memory;

[0083] The processor calls the executable program code stored in the memory to execute the steps of the ViT-based unsupervised domain adaptation image classification method described above.

[0084] In another embodiment, a computer-readable storage medium is provided that stores computer instructions, which, when invoked, are used to perform the steps of any of the above-described ViT-based unsupervised domain adaptation image classification methods.

[0085] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0086] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A system that specifies functions in one or more boxes.

[0087] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction set implemented in a process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0088] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0089] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0090] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An unsupervised domain-adaptive image classification method based on ViT, characterized in that, It includes the following steps: Step 1: Introduce a domain-based cross-attention module based on global correspondence into the original ViT to interact with sample features from different domains at the domain level and capture the global correspondence between cross-domain sample features; obtain sample-level global features based on the captured global correspondence between cross-domain sample features, input the sample-level global features into the original ViT, and the original ViT outputs multi-level global features. The global correspondence between cross-domain sample features is represented by an attention score matrix. The attention score matrix is ​​calculated based on the query generated by the source domain sample features and the key value generated by the target domain sample features. The target domain features are reconstructed based on the attention score matrix. The reconstructed target domain features continue to use the labels of the source domain features and are used as new source domain features, i.e., sample-level global features. The new source domain features are input into the original ViT, and the original ViT outputs multi-level global features. Step 2: In the original ViT, a convolution-transformer feature interaction module is introduced to input multi-level global features into the convolution-transformer feature interaction module to obtain local features. The obtained local features are then interacted with the multi-level global features and superimposed to obtain fused features. The process of obtaining local features is as follows: First, the input image is divided into several image blocks, and the flattened feature sequence of ViT is projected back into the 2D image space using the orthographic projection operator; Then, the two-dimensional convolution operator is repeatedly applied to the local blocks in the flattened feature sequence of ViT projected back to 2D space to generate local convolution features. For the obtained local convolutional features, the inverse projection operator is used to project them back to the original dimension in the feature space to obtain new features, which are local features. The process of obtaining fused features is as follows: After obtaining local features, the local features and multi-level global features are first interacted using a self-attention mechanism, and then the local features and multi-level global features are superimposed using residual connections to obtain fused features. Step 3: After introducing the domain cross-attention module and the convolution-transformer feature interaction module, the original ViT is transformed into a fused ViT. The fused ViT is then iteratively trained using adversarial training to obtain the trained fused ViT. Step four: Use the trained fusion ViT to perform the image classification task.

2. The unsupervised domain-adaptive image classification method based on ViT as described in claim 1, characterized in that, In a two-dimensional convolution operator, the kernel size is k. Zero padding of size (k-1) / 2 is added around the perimeter of the region covered by the kernel to keep the size of the local convolution feature unchanged.

3. The unsupervised domain-adaptive image classification method based on ViT as described in claim 1, characterized in that, Before step three is executed, the merged ViT is divided into several stages, with each stage handling patches of different sizes.

4. The unsupervised domain-adaptive image classification method based on ViT as described in claim 1, characterized in that, The loss function for adversarial training is as follows: ; Among them, the i-th sample and its corresponding one-hot tag Represents a labeled source set Using the j-th sample Represents the unlabeled target set , and These represent the sample sizes in the source and target domains, respectively. yes and union, Indicates a field label, The model's prediction results for the input samples. As a weighting factor, For feature extraction networks, For classifiers, It is the cross-entropy loss on labeled source domain samples. It is the classic domain alignment loss based on domain adversarial learning: ; in, Represents the i-th sample in the source domain. This represents the i-th sample in the target domain. This represents the i-th input sample. This represents the true label of the i-th sample in the source domain.

5. An unsupervised domain-adaptive image classification system based on ViT, characterized in that, The system includes the ViT architecture and: The domain cross-attention module is used to interact with sample features between different domains at the domain level, capture the global correspondence between cross-domain sample features; obtain sample-level global features based on the captured global correspondence between cross-domain sample features, input the sample-level global features into the original ViT, and the original ViT outputs multi-level global features. The global correspondence between cross-domain sample features is represented by an attention score matrix. The attention score matrix is ​​calculated based on the query generated by the source domain sample features and the key value generated by the target domain sample features. The target domain features are reconstructed based on the attention score matrix. The reconstructed target domain features continue to use the labels of the source domain features and are used as new source domain features, i.e., sample-level global features. The new source domain features are input into the original ViT, and the original ViT outputs multi-level global features. The convolution-transformer feature interaction module inputs multi-level global features to obtain local features. These local features are then interacted with the multi-level global features and superimposed to obtain fused features. The process of obtaining local features is as follows: First, the input image is divided into several image blocks, and the flattened feature sequence of ViT is projected back into the 2D image space using the orthographic projection operator; Then, the two-dimensional convolution operator is repeatedly applied to the local blocks in the flattened feature sequence of ViT projected back to 2D space to generate local convolution features. For the obtained local convolutional features, the inverse projection operator is used to project them back to the original dimension in the feature space to obtain new features, which are local features. The process of obtaining fused features is as follows: After obtaining local features, the local features and multi-level global features are first interacted using a self-attention mechanism, and then the local features and multi-level global features are superimposed using residual connections to obtain fused features. The training module integrates the domain cross-attention module and the convolution-transformer feature interaction module into ViT to form a fused ViT. Iterative training of the fused ViT is carried out using adversarial training to obtain a trained fused ViT. The execution module uses the trained fusion ViT to perform image classification tasks.

6. An electronic device, characterized in that, The device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the steps of the ViT-based unsupervised domain adaptation image classification method as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, which, when invoked, are used to perform the steps of the ViT-based unsupervised domain adaptation image classification method as described in any one of claims 1-4.