Domain Adaptation Method, Device, Terminal and Medium for Double-Layer Vision-Based Model
By adding two-layer visual condition flags to the source domain model and updating its parameters in real time, the problem that the passive domain adaptive method cannot update the model in real time during testing is solved, and the model's update efficiency and cross-domain robustness are improved.
Patent Information
- Application Number
- CN202410359260.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-03-27
AI Technical Summary
In the prior art, the passive field adaptive method cannot update the model online in real time during testing, and the update efficiency is low.
Add two-layer visual condition markers to the source domain model, which are used to learn long-term changes in domain-specific features and local changes in sample instance-specific features of domain-specific features, and update the parameters of the double-layer visual condition markers and normalization layer through backpropagation, and update the target domain model in real time.
Real-time online model update during testing is realized, the model update efficiency in actual test scenarios is improved, and the model's robustness in a cross-domain environment is enhanced.
Smart Images

Figure CN118212461B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transfer learning, and particularly to a domain adaptation method, device, terminal and medium for a model based on double-layer vision. Background Art
[0002] Although deep neural networks have achieved great success in various computer vision tasks, they come at the cost of a large amount of labeled training data sets and a large amount of computing resources. When the training samples and test samples come from different environments, their performance often drops significantly, which is usually referred to as the domain shift problem. Domain adaptation during testing is a set of transfer learning methods aimed at migrating the network model trained on a labeled data set (source domain) to a new unlabeled data set (target domain), and performing sequence analysis on the input samples during the inference stage, which can effectively adapt the source model during the testing process to solve the cross-domain performance degradation problem of deep neural networks.
[0003] Currently, unsupervised domain adaptation is divided into two categories, namely the method of active domain adaptation and the method of passive domain adaptation. The active domain adaptation method mainly includes statistical distribution matching and generative adversarial network methods to minimize domain differences. However, the active domain adaptation method must rely on the labeled data in the source domain. In practical applications, due to data privacy issues or transmission difficulties, it is usually impossible to access the source domain samples. The passive domain adaptation method directly performs unsupervised transfer learning in the target domain without the need for labeled data in the source domain. However, the passive domain adaptation method needs to access the entire test set data and requires multiple rounds of training, and cannot update the model online in real time during testing, resulting in low update efficiency in actual test scenarios.
[0004] Therefore, there are defects in the prior art and it needs to be improved and developed. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a domain adaptation method, device, terminal and medium for a model based on double-layer vision in view of the above-mentioned defects of the prior art, aiming to solve the problem that the passive domain adaptation method in the prior art cannot update the model online in real time during testing and has low update efficiency in actual test scenarios.
[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0007] A domain adaptation method for a model based on double-layer vision, wherein the method includes:
[0008] Obtain a target domain image, and input the target domain image into a pre-trained source domain model. A double-layer visual conditional flag is added to the source domain model, and the double-layer visual conditional flag is used to learn the long-term changes of domain-specific features and the local changes of sample instance-specific features of domain shift;
[0009] Use the embedding layer in the source domain model to convert the target domain image into an image patch vector, input the image patch vector and the double-layer visual conditional flag into the encoder of the source domain model, and update the parameters of the double-layer visual conditional flag and the normalization layer of the source domain model through backpropagation to obtain a target domain model;
[0010] Input the target domain image into the target domain model to obtain a corresponding image classification result.
[0011] In one implementation, the double-layer visual conditional flag includes: a domain-specific flag and a sample instance-specific flag. The domain-specific flag is used to learn the long-term changes of domain-specific features, and the sample instance-specific flag is used to learn the local changes of sample instance-specific features of domain shift;
[0012] The initial parameter of the domain-specific flag is the class flag pre-trained in the source domain model, and the initial parameter of the sample instance-specific flag is a zero vector.
[0013] In one implementation, obtaining a target domain image and inputting the target domain image into a pre-trained source domain model includes:
[0014] Obtain a target domain dataset for model domain adaptation, and the target domain dataset includes a number of target domain images;
[0015] Input all the target domain images in the target domain dataset into the pre-trained source domain model in batches.
[0016] In one implementation, using the embedding layer in the source domain model to convert the target domain image into an image patch vector, inputting the image patch vector and the double-layer visual conditional flag into the encoder of the source domain model, and updating the parameters of the double-layer visual conditional flag and the normalization layer of the source domain model through backpropagation to obtain a target domain model includes:
[0017] Use the embedding layer in the source domain model to separately slice all the target domain images in the current batch into a number of image patches, and convert each image patch into an image patch vector;
[0018] Input the image patch vector, domain-specific flag, and sample instance-specific flag into the encoder of the source domain model, and use a preset information entropy loss function to backpropagate and update the parameters of the domain-specific flag, sample instance-specific flag, and the normalization layer of the source domain model to obtain the target domain model.
[0019] In one implementation, several target domain images in each batch are simultaneously input into the pre-trained source domain model, and the source domain model used in the current batch is the target domain model updated from the previous batch.
[0020] In one implementation, when updating the parameters of the domain-specific flag, sample instance-specific flag, and the normalization layer of the source domain model, the gradient descent method is used; before the input of the target domain images in each batch, the parameters and gradients of the sample instance-specific flag in the target domain model obtained from the previous batch are set to zero.
[0021] In one implementation, input the target domain images into the target domain model to obtain the corresponding image classification results, including:
[0022] Input all the target domain images in the current batch into the target domain model to obtain the image classification results corresponding to each target domain image.
[0023] The present invention discloses a model domain adaptation device based on double-layer vision, wherein the device includes:
[0024] An input module, configured to obtain target domain images and input the target domain images into a pre-trained source domain model, and a double-layer vision condition flag is added to the source domain model, and the double-layer vision condition flag is used to learn the long-term changes of domain-specific features and the local changes of sample instance-specific features of domain shifts;
[0025] An update module, configured to use the embedding layer in the source domain model to convert the target domain images into image patch vectors, input the image patch vectors and the double-layer vision condition flag into the encoder of the source domain model, and backpropagate to update the parameters of the double-layer vision condition flag and the normalization layer of the source domain model to obtain the target domain model;
[0026] A classification module, configured to input the target domain images into the target domain model to obtain the corresponding image classification results.
[0027] The present invention discloses a terminal, which includes: a memory, a processor, and a model domain adaptation program based on double-layer vision stored on the memory and executable on the processor. When the model domain adaptation program based on double-layer vision is executed by the processor, the steps of the model domain adaptation method based on double-layer vision described above are implemented.
[0028] The present invention discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the above-mentioned domain adaptation method for a double-layer vision-based model.
[0029] The domain adaptation method, device, terminal and medium for a double-layer vision-based model provided by the present invention, the method includes: acquiring a target domain image, inputting the target domain image into a pre-trained source domain model, and a double-layer vision condition flag is added to the source domain model, and the double-layer vision condition flag is used to learn the long-term change of domain-specific features and the local change of sample instance-specific features of domain shift; converting the target domain image into an image patch vector by using an embedding layer in the source domain model, inputting the image patch vector and the double-layer vision condition flag into an encoder of the source domain model, and updating the parameters of the double-layer vision condition flag and the normalization layer of the source domain model through backpropagation to obtain a target domain model; inputting the target domain image into the target domain model to obtain a corresponding image classification result. By adding a double-layer vision condition flag to the source domain model and being able to update the parameters of the double-layer vision condition flag and the normalization layer of the source domain model in real-time through backpropagation, the present invention realizes real-time online model update and improves the model update efficiency in actual test scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flowchart of a preferred embodiment of the domain adaptation method for a double-layer vision-based model in the present invention;
[0031] Figure 2 is a logical schematic block diagram of a preferred embodiment of the domain adaptation method for a double-layer vision-based model in the present invention;
[0032] Figure 3 is a functional principle block diagram of a preferred embodiment of the domain adaptation device for a double-layer vision-based model in the present invention;
[0033] Figure 4 is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0035] In the transfer learning image classification task, when the model trained in the source domain is transferred to the target domain, its performance often deteriorates significantly. Most of the existing domain adaptation methods in the field of transfer learning require access to the training data of the source domain, which is not applicable in some scenarios where data transmission is difficult or privacy is sensitive. To address such data transmission or privacy issues, source-free domain adaptation that does not require access to the source domain data has been proposed, and current source-free domain adaptation methods require access to the entire test set data and multiple rounds of training. However, the embodiments of this application propose complete test-time adaptation, which only requires access to the real-time stream of test samples and can dynamically adapt the source model during the test process. Since the key task of test-time adaptation is to learn domain-specific information and separate the domain-specific information from the input sample features. In this work, a vision Transformer (ViT) can be used as the backbone encoder for the image classification task. In ViT, the class token is trained to capture the semantic information of the source image. Due to the distribution shift, this pre-trained token may not generalize to the target domain image because the prior knowledge about the semantics embedded in the class token is source-domain specific, and the pre-trained class token has no prior knowledge about the domain shift it encounters during testing.
[0036] The present invention discovers that during the test-time adaptation process, the class token of the first layer encoder can be learned to learn the domain shift characteristics, which is called the Visual Conditioning Token (VCT). Once the visual conditioning token is successfully learned, an adjustment operation can be performed on the input image to correct the image perturbation caused by the domain shift. As the domain shift perturbation or corruption is removed from the image features, the network model will be more robust to the domain shift, thus significantly improving the generalization performance. At the domain level, different distribution changes have different impacts on the overall model performance of the test dataset. In addition, at each individual image level, the distribution shift has different impacts on the performance of different images. Therefore, to successfully learn the visual conditioning token, the present invention proposes a two-layer learning method to characterize the long-term changes of domain-specific features and the local changes of instance-specific features of domain shift.
[0037] The objective of the present invention is to solve the problem of the decrease in the accuracy of the test set caused by domain drift in transfer learning by learning a complete test-time adaptation model through two-layer visual conditioning token learning, learning the long-term changes of domain-specific features and the local changes of instance-specific features of domain shift of the class token of the first layer encoder, and effectively adapting the source model during the test process to solve the problem of cross-domain performance degradation of deep neural networks.
[0038] Please refer to Figure 1 , Figure 1 , which is a flowchart of the domain adaptation method for the double-layer vision-based model in the present invention. As Figure 1 shown, the domain adaptation method for the double-layer vision-based model described in the embodiments of the present invention includes:
[0039] Step S100: Obtain a target domain image, and input the target domain image into a pre-trained source domain model. A double-layer vision conditioning flag is added to the source domain model, and the double-layer vision conditioning flag is used to learn the long-term changes of domain-specific features and the local changes of sample instance-specific features of domain shift.
[0040] Specifically, since class tokens can learn domain-specific information in the Transformer model, the present invention proposes a new type of double-layer vision conditioning token learning fully test-time adaptation model to solve the problem of the decrease in the accuracy of the test set caused by domain drift in transfer learning. In image classification based on the Transformer model, the class tokens of the first layer encoder can be learned, and sequence analysis is performed on the target domain image during the inference stage to capture the domain-specific features of the target domain image during test-time adaptation. The double-layer learning method proposed by the present invention can learn the long-term changes of domain-specific features and at the same time adapt to the local changes of instance-specific features, and can effectively adapt the source model during the test process to solve the problem of cross-domain performance degradation of deep neural networks.
[0041] In the embodiments of the present application, the double-layer vision conditioning flag includes: a domain-specific flag and a sample instance-specific flag. The domain-specific flag is used to learn the long-term changes of domain-specific features, and the sample instance-specific flag is used to learn the local changes of sample instance-specific features of domain shift. The initialization parameter of the domain-specific flag is the class token pre-trained in the source domain model, and the initialization parameter of the sample instance-specific flag is a zero vector.
[0042] Specifically, first train a source domain model in the source domain, and then perform unsupervised domain adaptation in the target domain. The source domain model can use any pre-trained Transformer model, and initialize the parameter θ L of the domain-specific flag as the pre-trained class token, and the class token is the class token of the first layer encoder, and initialize the parameter θ S of the sample instance-specific flag as a 0 vector. The formula is as follows:
[0043]
[0044]
[0045] Among them, η l represents the learning rate of the domain - specific flag, η s represents the learning rate of the sample - instance - specific flag, represents the gradient of the domain - specific flag, represents the gradient of the sample - instance - specific flag, θ t represents the weight parameter of the entire model at time t, and x represents the input target - domain image.
[0046] In the embodiment of the present application, the step S100 specifically includes:
[0047] Step S110, obtaining a target - domain data set for model domain adaptation, where the target - domain data set includes a number of target - domain images;
[0048] Step S120, inputting all the target - domain images in the target - domain data set into the pre - trained source - domain model in batches.
[0049] Specifically, in the test and inference stage of the present invention, domain adaptation can be performed by only accessing the current batch of data. When the pre - trained Transformer model is migrated to the downstream test scenario, inputting the pre - trained source - domain model in batches can update the domain - specific flag, sample - instance - specific flag, and normalization layer in real time, obtaining the target - domain model, which improves the update efficiency. The existing technical solutions do not utilize the long - term changes of domain - specific features learned by class flags and the local changes of sample - instance - specific features of domain shifts, and need to access the entire target - domain data and perform multiple trainings to achieve model update, with low update efficiency.
[0050] As Figure 1 shown, the model domain adaptation method based on double - layer vision in this embodiment further includes:
[0051] Step S200, using the embedding layer in the source - domain model to convert the target - domain image into an image - block vector, inputting the image - block vector and the double - layer vision condition flag into the encoder of the source - domain model, and backpropagating to update the parameters of the double - layer vision condition flag and the normalization layer of the source - domain model to obtain the target - domain model.
[0052] By backpropagating to update the double - layer vision condition flag and the normalization layer of the source - domain model in the embodiment of the present application, the performance of domain adaptation during testing can be significantly improved, and it is much better than the existing methods.
[0053] In the embodiment of the present application, the step S200 specifically includes:
[0054] Step S210, using the embedding layer in the source - domain model to respectively cut each of the target - domain images in the current batch into a number of image blocks, and converting each of the image blocks into an image - block vector;
[0055] Step S220: Input the image patch vector, domain-specific flag, and sample instance-specific flag into the encoder of the source domain model, and use a preset information entropy loss function to backpropagate and update the parameters of the domain-specific flag, sample instance-specific flag, and the normalization layer of the source domain model to obtain the target domain model.
[0056] Please refer to Figure 2 , the target domain image is segmented into a fixed number of image patches, and each image patch is usually a small square region. Each image patch is converted into a vector representation through an embedding layer, usually by linearly projecting the pixel values of the image patch into a vector of a fixed dimension. Position encoding is added to the embedding vector of each image patch to capture the position information of the image patch in the original image. In the first layer of the encoder, in addition to the embedding vector of the image patch, the present invention also adds a domain-specific flag and a sample instance-specific flag. The image patch vector, domain-specific flag, and sample instance-specific flag after passing through the embedding layer and position encoding are input into the Transformer encoder. The Transformer encoder uses the self-attention mechanism to capture the global relationships and dependencies between image patches, thereby encoding the input sequence. The ViT model usually contains multiple Transformer encoder layers, and each layer processes the input and passes it to the next layer to gradually extract and integrate image information. The output of the last Transformer encoder is fed into a multi-layer perceptron (MLP) for classification tasks.
[0057] After the present invention inputs the target domain image of the current batch, it uses the information entropy loss function of the output probability distribution to backpropagate and update the domain-specific flag, sample instance-specific flag, and the normalization layer (LayerNormalization layer) of the source domain model, thereby achieving access to the data of the current batch and dynamically fine-tuning the model, so as to improve the robustness of the model across domains.
[0058] Among them, the information entropy loss function is:
[0059]
[0060] Among them, E0 is the information entropy threshold for filtering out samples with excessive noise; E is the information entropy function; B j represents a batch of test samples.
[0061] In addition, when there are multiple images in the current batch, the information entropy loss function takes the average of the information entropy loss functions corresponding to each image.
[0062] In one embodiment of the present application, several target domain images of each batch are simultaneously input into a pre-trained source domain model, and the source domain model used in the current batch is the target domain model updated from the previous batch.
[0063] Different from existing domain adaptation methods, the present invention does not require accessing source domain data or using the entire target domain test set data for multiple iterative trainings. It only needs to access the current batch of data during test inference for domain adaptation, and dynamically fine-tune the model to improve the robustness of the model across domains.
[0064] The present invention uses a two-layer visual conditional flag learning fully test-time adaptation model to learn the class flags of the first-layer encoder to learn the long-term changes of domain-specific features and the local changes of instance-specific features of domain shifts, so as to solve the problem of the decrease in the accuracy of the test set caused by domain drift in transfer learning, and effectively adapt the source model during the test process to solve the problem of cross-domain performance degradation of deep neural networks.
[0065] In one embodiment of the present application, when updating the parameters of the domain-specific flag, the instance-specific flag, and the normalization layer of the source domain model, the gradient descent method is adopted; before the target domain images of each batch are input, the parameters and gradients of the instance-specific flag in the target domain model obtained in the previous batch are both set to zero.
[0066] Specifically, when testing, a batch of test samples B j is given, and the parameters θ L of the domain-specific token and the parameters θ S of the instance-specific token and the parameters of the normalization layer are updated respectively using gradient descent. After the update, the prediction results of the current batch are output. Among them, the parameters θ S and gradients are both set to zero after the prediction results of each batch. That is to say, between different batches, the domain-specific flag and the normalization layer are cumulatively updated, while the instance-specific flag is reset to 0 after each batch update. This is because the features between different batches are different.
[0067] As Figure 1 shown, the model domain adaptation method based on two-layer vision described in this embodiment further includes:
[0068] Step S300: Input the target domain image into the target domain model to obtain the corresponding image classification result.
[0069] In one embodiment of the present application, step S300 specifically includes: inputting all the target domain images in the current batch into the target domain model to obtain the image classification results corresponding to each target domain image.
[0070] Specifically, the present invention only needs to access the data of one batch to update the source domain model, achieving real-time update. There is no need to use all the test data for multiple iterative trainings, improving the update efficiency.
[0071] In one embodiment, as Figure 3 shown, based on the above-mentioned model domain adaptation method based on double-layer vision, the present invention also correspondingly provides a model domain adaptation device based on double-layer vision, including:
[0072] An input module 100, configured to obtain target domain images and input the target domain images into a pre-trained source domain model, where a double-layer vision condition flag is added to the source domain model, and the double-layer vision condition flag is used to learn the long-term changes of domain-specific features and the local changes of sample instance-specific features of domain offsets;
[0073] An update module 200, configured to convert the target domain images into image block vectors by using the embedding layer in the source domain model, input the image block vectors and the double-layer vision condition flag into the encoder of the source domain model, and update the parameters of the double-layer vision condition flag and the normalization layer of the source domain model through backpropagation to obtain a target domain model;
[0074] A classification module 300, configured to input the target domain images into the target domain model to obtain corresponding image classification results.
[0075] Figure 4 It is a schematic structural diagram of a terminal provided in an embodiment of the present application. The terminal may include:
[0076] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.
[0077] When the processor 502 executes the program, it implements the model domain adaptation method based on double-layer vision provided in the above embodiment.
[0078] Furthermore, the terminal further includes:
[0079] A communication interface 503, configured for communication between the memory 501 and the processor 502.
[0080] The memory 501 is used to store a computer program executable on the processor 502.
[0081] The memory 501 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0082] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0083] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.
[0084] The processor 502 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0085] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned method for domain adaptation of a model based on double-layer vision is implemented.
[0086] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0087] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0088] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application pertain.
[0089] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can read and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or N wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0090] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0091] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiments.
[0092] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0093] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
[0094] In summary, the present invention discloses a method, device, terminal, and medium for domain adaptation of a model based on double-layer vision. The method includes: obtaining a target domain image, inputting the target domain image into a pre-trained source domain model, where a double-layer vision condition flag is added to the source domain model, and the double-layer vision condition flag is used to learn the long-term changes of domain-specific features and the local changes of sample instance-specific features of domain shifts; using the embedding layer in the source domain model to convert the target domain image into an image patch vector, inputting the image patch vector and the double-layer vision condition flag into the encoder of the source domain model, and backpropagating to update the parameters of the double-layer vision condition flag and the normalization layer of the source domain model to obtain a target domain model; inputting the target domain image into the target domain model to obtain a corresponding image classification result. By adding a double-layer vision condition flag to the source domain model and being able to update the parameters of the double-layer vision condition flag and the normalization layer of the source domain model in real-time through backpropagation, the present invention realizes real-time online model update and improves the model update efficiency in actual test scenarios.
[0095] It should be understood that the application of the present invention is not limited to the above examples. Those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A model domain adaptation method based on double-layer vision, characterized in that: The method comprises: Acquire a target domain image, and input the target domain image into a pre-trained source domain model, wherein a double-layer visual condition flag is added to the source domain model, and the double-layer visual condition flag is used to learn long-term changes in domain-specific features and local changes in sample instance-specific features of domain shift; The target domain image is converted into an image block vector by using an embedding layer in the source domain model, the image block vector and the two-layer visual condition flag are input into an encoder of the source domain model, and the two-layer visual condition flag and the parameters of the normalization layer of the source domain model are updated by back propagation to obtain a target domain model; Inputting the target domain image into the target domain model to obtain a corresponding image classification result; The dual-layer visual condition mark includes: a domain-specific mark and a sample instance-specific mark, wherein the domain-specific mark is used to learn the long-term change of the domain-specific feature, and the sample instance-specific mark is used to learn the local change of the sample instance-specific feature of the domain shift; the initialization parameter of the domain-specific mark is the class mark pre-trained in the source domain model, and the initialization parameter of the sample instance-specific mark is a zero vector; the class mark is the class mark of the first layer encoder; In which, the gradient descent method is used when updating the parameters of the domain-specific flag, the sample instance-specific flag and the normalization layer of the source domain model; before each batch of target domain images is input, the parameters and gradients of the sample instance-specific flag in the target domain model obtained in the previous batch are set to zero.
2. The model domain adaptation method based on dual-layer vision according to claim 1, characterized in that: Obtaining a target domain image and inputting the target domain image into a pre-trained source domain model includes: Acquire a target domain dataset for model domain adaptation, wherein the target domain dataset includes a plurality of target domain images; All target domain images in the target domain dataset are input into the pre-trained source domain model in batches.
3. The model domain adaptation method based on double-layer vision according to claim 2 is characterized in that: The target domain image is converted into an image block vector by using an embedding layer in the source domain model, the image block vector and the two-layer visual condition flag are input into an encoder of the source domain model, and the two-layer visual condition flag and the parameters of the normalization layer of the source domain model are updated by back propagation to obtain a target domain model, including: Using the embedding layer in the source domain model, all the target domain images in the current batch are divided into a number of image blocks, and each of the image blocks is converted into an image block vector; The image block vector, domain-specific flag and sample instance-specific flag are input into the encoder of the source domain model, and the domain-specific flag, sample instance-specific flag and parameters of the normalization layer of the source domain model are updated by back-propagation using a preset information entropy loss function to obtain a target domain model.
4. The model domain adaptation method based on dual-layer vision according to claim 3 is characterized in that: Several target domain images in each batch are simultaneously input into the pre-trained source domain model, and the source domain model used in the current batch is the target domain model updated in the previous batch.
5. The model domain adaptation method based on dual-layer vision according to claim 2, characterized in that: Inputting the target domain image into the target domain model to obtain a corresponding image classification result includes: All the target domain images in the current batch are input into the target domain model to obtain image classification results corresponding to each target domain image.
6. A model domain adaptation device based on double-layer vision, characterized in that: The device comprises: An input module, used for acquiring a target domain image, and inputting the target domain image into a pre-trained source domain model, wherein a double-layer visual condition flag is added to the source domain model, and the double-layer visual condition flag is used for learning long-term changes of domain-specific features and local changes of sample instance-specific features of domain shift; An updating module, configured to convert the target domain image into an image block vector by using an embedding layer in the source domain model, input the image block vector and the two-layer visual condition flag into an encoder of the source domain model, and back-propagate and update the two-layer visual condition flag and a parameter of a normalization layer of the source domain model to obtain a target domain model; A classification module, used for inputting the target domain image into the target domain model to obtain a corresponding image classification result; The dual-layer visual condition mark includes: a domain-specific mark and a sample instance-specific mark, wherein the domain-specific mark is used to learn the long-term change of the domain-specific feature, and the sample instance-specific mark is used to learn the local change of the sample instance-specific feature of the domain shift; the initialization parameter of the domain-specific mark is the class mark pre-trained in the source domain model, and the initialization parameter of the sample instance-specific mark is a zero vector; the class mark is the class mark of the first layer encoder; In which, the gradient descent method is used when updating the parameters of the domain-specific flag, the sample instance-specific flag and the normalization layer of the source domain model; before each batch of target domain images is input, the parameters and gradients of the sample instance-specific flag in the target domain model obtained in the previous batch are set to zero.
7. A terminal, characterized in that: include: A memory, a processor, and a model domain adaptation program based on dual-layer vision stored in the memory and executable on the processor, wherein the model domain adaptation program based on dual-layer vision, when executed by the processor, implements the steps of the model domain adaptation method based on dual-layer vision as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed to implement the steps of the model domain adaptation method based on dual-layer vision as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Passive field adaptive target detection method
CN112861616A
Domain generalization image classification method and system based on feature adjustment, terminal and medium
CN117115567A