Unsupervised Domain Adaptation Method Based on Bidirectional Matching
By adopting an unsupervised method of bidirectional matching in domain adaptation, the two models learn from each other and through consistency regularization and threshold setting, the problem of cumbersome domain adaptation process in the prior art is solved, and the domain adaptation effect with high accuracy and flexibility is achieved.
Patent Information
- Application Number
- CN202211388976.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-11-03
AI Technical Summary
In the prior art, if there is already one type of data set and you want to identify another form, you need to prepare another data set, which is cumbersome and difficult to achieve domain adaptation.
An unsupervised domain adaptation method based on bidirectional matching is proposed. Through learning from two models, the overall accuracy of the model is improved, the model is overfitted through consistency regularization, and the threshold is set during the inference process to ensure the reliability of the results.
Transfer learning from the source domain to the target domain is realized, learning from only source domain samples and a small number of target domain samples can be performed, and target domain samples can be classified, which improves the accuracy of the model and the flexibility of application.
Smart Images

Figure CN115620091B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern recognition and computer vision, and particularly relates to an unsupervised domain adaptation method based on bidirectional matching. Background Art
[0002] Currently, most neural networks rely on training on large datasets to obtain a model for performing various tasks. However, in some tasks, the object to be recognized remains the same, but its manifestation form changes. For example, a real car and a car in a painting are both cars but have different forms. At this time, if there is a dataset of one type and one wants to recognize another form, a new dataset needs to be prepared, which is a cumbersome process.
[0003] Therefore, the domain adaptation method came into being. Domain adaptation can predict the target domain dataset with only the source domain dataset and a small amount of target domain data, thereby reducing the intermediate dataset preparation process and increasing the flexibility of applications. Summary of the Invention
[0004] To solve the problem of domain adaptation, the present invention proposes an unsupervised domain adaptation method based on bidirectional matching, which improves the overall accuracy of the model by the way of two models learning from each other, prevents overfitting during the model convergence process through consistency regularization, and ensures the reliability of the results by setting a threshold during the inference process, so that the entire network result can achieve a more accurate domain adaptation effect.
[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0006] An unsupervised domain adaptation method based on bidirectional matching, characterized by including the following steps:
[0007] Step S1: Mix the image data of the source domain and the target domain, and perform operations including random rotation on the mixed images for data augmentation;
[0008] Step S2: Train the neural network by a machine learning method based on confidence; during the training process, ensure the stability of model convergence through consistency regularization;
[0009] Step S3: Fine-tune the target domain model to ensure the reliability of the target domain model;
[0010] Step S4: Perform model inference, and ensure the stability of the inference result by setting a threshold.
[0011] Further, step S1 specifically includes the following steps:
[0012] Step S11: Obtain the publicly available domain adaptation dataset, the office-31 dataset, and obtain the annotations of the training data;
[0013] Step S12: Use two fixed mixing ratio coefficients λ sd , λ td to perform data mixing; Given a pair of input sample images and a label corresponding to them in the source and target domains: Mix the sample images and labels through the following equation: where λ sd represents the mixing ratio coefficient of the sample image from the source domain, λ td represents the mixing ratio coefficient of the sample image from the target domain, and λ sd + λ td = 1, represents the sample image from the source domain, represents the label of the sample image from the source domain, represents the sample image from the target domain, represents the estimated value of the label of the sample image from the target domain, represents the sample image after mixing, represents the label after mixing, where s represents from the source domain, t represents from the target domain, and i represents which batch process is currently in during the training process;
[0014] Step S13: Perform random transformation on the sample image obtained after mixing for data augmentation. The specific operations are as follows: Randomly select 0 to 3 items from (1), (2), (3) for operations. (1) Randomly rotate the image around the center point, and the rotation angle is randomly selected from 0° to 360°. The obtained sample image is cropped with the original image size. (2) Randomly select one of horizontal flipping and vertical flipping for the sample image to perform. (3) Randomly scale the sample image around the center point, and the scaling factor is randomly selected from 0.5 to 1.2. The obtained sample image is cropped with the original image size.
[0015] Furthermore, Step S2 includes the following steps:
[0016] Step S21: Input the preprocessed sample images and labels obtained in Step S1 into the source domain model and the target domain model respectively. Among them, both the source domain model and the target domain model are constructed based on Resnet50. When inputting the preprocessed sample images and labels into the source domain model, ensure that λ sd > λ td during the mixing process in Step S12. When inputting the preprocessed sample images and labels into the target domain model, ensure that λ sd < λ td during the mixing process in Step S12. The probability of the class C s with the highest probability output by the source domain model is The probability of the class with the second highest probability is The class C with the highest probability output by the target domain model s The probability of The probability of the class with the second highest probability is
[0017] Step S22: Use the class C s as a label and input it into the target domain model together with the corresponding sample images in the next training batch to obtain the probability on the class C s The maximum probability among all classes Increase the probability of class C in the target domain model by minimizing the objective function, and the objective function is: s where L represents the loss function on the target domain, B represents the size of each batch, max represents the maximum value function, exp represents the exponential function, and the probability of the class C with the highest probability output by the source domain model t is s The probability of The probability of the class with the second highest probability is i represents which batch process is currently in the training process;
[0018] Step S23: If where p border is the set threshold, is the maximum value of the probabilities of all classes output by the target domain model, then increase the probabilities of all classes in the source domain model except for the class C with the highest probability t The specific implementation method is to minimize the objective function: where is the probability of the source domain model on the class C t The probability of the class C with the highest probability output by the source domain model s is L s represents the loss function on the source domain, and the probability of the class C with the highest probability output by the target domain model t is The probability of the class with the second highest probability is B represents the size of each batch, max represents the maximum value function, exp represents the exponential function, and i represents which batch process is currently in the training process.
[0019] Step S24: To ensure that the source domain model and the target domain model do not overfit and enhance the convergence stability, add a regularization term to the objective function of the source domain model during training, and the regularization term is defined as: where Denote the regularization term of the objective function on the source domain; add a regularization term to the objective function of the target domain model, and the regularization term is defined as: where denotes the regularization term of the objective function on the target domain.
[0020] Furthermore, step S3 specifically includes the following steps:
[0021] Step S31: After the training batch is greater than b batch , until the training ends, fine-tune the target domain model, where b batch is the batch number set manually. The specific operation is as follows: Duplicate a single sample from the target domain 10 times, and each copy is randomly preprocessed according to the preprocessing method in step S13 to obtain 10 preprocessed samples. Input these 10 preprocessed samples into the target domain model respectively to obtain their outputs. If the number of outputs with the same result is greater than b sc , where b sc is the threshold set manually, then minimize the cross-entropy loss with this same result as the label. If the number of outputs with the same result is less than b sc , then maximize the cross-entropy loss with the most same results as the label.
[0022] Step S32: Since during the training process of the source domain model, the mixed samples and labels are always used as inputs, the source domain data features are not extracted well enough. Therefore, input the original samples and labels in the source domain into the source domain model to fine-tune the source domain model, and the most traditional way of minimizing the cross-entropy loss is used during the training process.
[0023] Furthermore, step S4 specifically includes the following steps:
[0024] Step S41: During inference, input the pictures to be classified into the source domain model and the target domain model respectively. The probability that the source domain model outputs on class k is The probability that the target domain model outputs on class k is where 0 < k < class + 1, class is the total number of classes, and calculate p k The class k corresponding to the maximum value 1 is the inference result;
[0025] Step S42: If the maximum value of p k among all classes is greater than the threshold Y, then keep the class k 1 as the inference result; if the maximum value of p k among all classes is less than the threshold Y, and the class k k with the second largest value of p 2 satisfies where is the probability output by the target domain model for class k 2 on, is the probability output by the target domain model for class k 1 on, then change the inference result to k 2 .
[0026] Compared with the prior art, the present invention and its preferred embodiments have the following beneficial effects:
[0027] 1. It can effectively perform transfer learning from the source domain to the target domain for the unsupervised domain adaptation problem, and can classify target domain samples by learning only from source domain samples and a small number of target domain samples.
[0028] 2. By using the source domain model and the target domain model to perform supervised learning on the source domain and the target domain respectively, the source domain model focuses on the data features of the source domain, and the target domain model focuses on the data features of the target domain. The overall accuracy of the model is improved by the way of mutual learning between the two models.
[0029] 3. By consistency regularization, overfitting during the model convergence process is prevented, and the stability of the two models with different focuses during convergence is ensured.
[0030] 4. By means of staged training, the model is fine-tuned near the end of training to ensure that the source domain model focuses on the source domain data and the target domain model focuses on the target domain data.
[0031] 5. By setting a threshold during the inference process, the reliability of the result is ensured. Since the target domain model focuses on the target domain features, when the threshold is not reached, the output of the target domain model is preferentially used to determine the classification result. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0033] Figure 1 is the flowchart of the method implementation of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0035] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0036] Note that the terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0037] As Figure 1 shown, the unsupervised domain adaptation method based on bidirectional matching provided by the embodiments of the present invention includes the following steps:
[0038] Step S1: Mix the image data of the source domain and the target domain, and perform operations including random rotation on the mixed images for data augmentation;
[0039] Step S2: Train the neural network by a machine learning method based on confidence; during the training process, ensure the stability of model convergence through consistency regularization;
[0040] Step S3: Fine-tune the target domain model to ensure the reliability of the target domain model;
[0041] Step S4: Perform model inference and ensure the stability of the inference result by setting a threshold.
[0042] Further, step S1 specifically includes the following steps:
[0043] Step S11: Obtain the publicly available domain adaptation dataset, the office-31 dataset, and obtain the annotations of the training data;
[0044] Step S12: Use two fixed mixing ratio coefficients λ sd , λ td to perform data mixing; given a pair of input sample pictures and a label corresponding to them in the source and target domains: Mix the sample pictures and labels through the following equation: where λ sd represents the mixing ratio coefficient of the sample picture from the source domain, λ td represents the mixing ratio coefficient of the sample picture from the target domain, and λ sd + λ td = 1, represents the sample image from the source domain, represents the label of the sample image from the source domain, represents the sample image from the target domain, represents the estimated value of the label of the sample image from the target domain, represents the sample image after mixing, Indicates the label after mixing, where s represents from the source domain, t represents from the target domain, and i represents which batch process is currently in during the training process;
[0045] Step S13: Randomly transform the obtained sample images after mixing for data augmentation. The specific operations are as follows: Randomly select 0 to 3 items from (1), (2), and (3) for operations. (1) Randomly rotate the image around the center point, and the rotation angle is randomly selected from 0° to 360°. The obtained sample image is cropped with the original image size. (2) Randomly select one of horizontal flipping and vertical flipping for the sample image to perform. (3) Randomly scale the sample image around the center point, and the scaling factor is randomly selected from 0.5 to 1.2. The obtained sample image is cropped with the original image size.
[0046] Further, step S2 includes the following steps:
[0047] Step S21: Input the preprocessed sample images and labels obtained in step S1 into the source domain model and the target domain model respectively. Among them, both the source domain model and the target domain model are constructed based on Resnet50. When inputting the preprocessed sample images and labels into the source domain model, ensure λ sd >λ td during the mixing process in step S12. When inputting the preprocessed sample images and labels into the target domain model, ensure λ sd <λ td during the mixing process in step S12. The probability of the class C s with the highest probability output by the source domain model is The probability of the class with the second highest probability is The probability of the class C s with the highest probability output by the target domain model is The probability of the class with the second highest probability is
[0048] Step S22: Use the class C s as the label and input it into the target domain model together with the corresponding sample images in the next training batch to obtain the probability s on the class C The maximum value of the probabilities among all classes Increase the probability of the class C s in the target domain model, which is achieved by minimizing the objective function. The objective function is: where L t represents the loss function on the target domain, B represents the size of each batch, max represents the maximum value function, exp represents the exponential function, and the probability of the class C s with the highest probability output by the source domain model is The probability of the second most likely class is i represents which batch process is currently in during the training process;
[0049] Step S23: If where p border is a set threshold, is the maximum value of the probabilities of all classes output by the target domain model, then increase the probabilities of all classes in the source domain model except for the class C with the highest probability t other than; The specific implementation method is to minimize the objective function: where is the probability of the source domain model for class C t The class C with the highest probability output by the source domain model s has a probability of L s represents the loss function on the source domain. The class C with the highest probability output by the target domain model t has a probability of The probability of the second most likely class is B represents the size of each batch, max represents the maximum value function, exp represents the exponential function, and i represents which batch process is currently in during the training process.
[0050] Step S24: To ensure that the source domain model and the target domain model do not overfit and enhance convergence stability, add a regularization term to the objective function of the source domain model during training. The regularization term is defined as: where represents the regularization term of the objective function on the source domain; add a regularization term to the objective function of the target domain model. The regularization term is defined as: where represents the regularization term of the objective function on the target domain.
[0051] Furthermore, in step S3, it specifically includes the following steps:
[0052] Step S31: After the training batch is greater than b batch and until the training ends, fine-tune the target domain model, where b batch is a manually set number of batches. The specific operation is as follows: Duplicate a single sample from the target domain 10 times, and perform random preprocessing on each copy according to the preprocessing method in step S13 to obtain 10 preprocessed samples. Input these 10 preprocessed samples into the target domain model respectively to obtain their outputs. If the number of identical outputs is greater than b sc , where b sc is a manually set threshold, then minimize the cross-entropy loss with this identical result as the label. If the number of identical outputs is less than bsc , the cross-entropy loss with the most same results as the label is maximized.
[0053] Step S32: Since the samples and labels after mixing are used as inputs during the training process of the source domain model, the source domain data features are not extracted well. Therefore, the original samples and labels in the source domain are input into the source domain model to fine-tune the source domain model, and the most traditional way of minimizing the cross-entropy loss is used for training during the training process.
[0054] Furthermore, in step S4, it specifically includes the following steps:
[0055] Step S41: During inference, the pictures to be classified are respectively input into the source domain model and the target domain model. The probability output by the source domain model on class k is The probability output by the target domain model on class k output is where 0 < k < class + 1, class is the total number of categories, and it is calculated for all classes p k The class k corresponding to the maximum value 1 is the inference result;
[0056] Step S42: If the maximum value of p k among all classes is greater than the threshold Y, then keep the class k 1 as the inference result; if the maximum value of p k among all classes is less than the threshold Y, and the class k k with the second largest value of p 2 among all classes satisfies where is the probability output by the target domain model on class k 2 , is the probability output by the target domain model on class k 1 , then change the inference result to k 2 .
[0057] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0058] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0059] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0060] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.
[0061] As mentioned above, it is only a preferred embodiment of the present invention, and it is not a limitation of the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.
[0062] This patent is not limited to the above best implementation manner. Anyone inspired by this patent can obtain various other forms of unsupervised domain adaptation methods based on bidirectional matching. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by this patent.
Claims
1. An unsupervised domain adaptation method based on bidirectional matching, characterized in that, it includes the following steps: Step S1: Mix the image data of the source domain and the target domain, and perform operations including random rotation on the mixed images for data augmentation; Step S2: Train the neural network through a machine learning method based on confidence; during the training process, ensure the stability of model convergence through consistency regularization; Step S3: Fine-tune the target domain model to ensure the reliability of the target domain model; Step S4: Perform model inference, and ensure the stability of the inference result by setting a threshold; Step S2 includes the following steps: Step S21: Input the preprocessed sample images and labels obtained in Step S1 into the source domain model and the target domain model respectively. Among them, both the source domain model and the target domain model are constructed based on Resnet50. When the preprocessed sample images and labels input into the source domain model are mixed in Step S12, ensure that λ sd >λ td , and when the preprocessed sample images and labels input into the target domain model are mixed in Step S12, ensure that λ sd <λ td . The probability of the class C s with the highest probability output by the source domain model is . The probability of the class with the second highest probability is . The probability of the class C s with the highest probability output by the target domain model is . The probability of the class with the second highest probability is Step S22: Take class C s as a label and input it into the target domain model together with the corresponding sample images in the next training batch to obtain the probability s on class C the maximum probability among all classes Increase the probability of class C s in the target domain model, which is achieved by minimizing the objective function. The objective function is: where L t represents the loss function on the target domain, B represents the size of each batch, max represents the maximum value function, exp represents the exponential function, and the probability of the class C s with the highest probability output by the source domain model is the probability of the class with the second highest probability is i represents which batch process is currently in during the training process; Step S23: If where p border is a set threshold, is the maximum value of the probabilities of all classes output by the target domain model, then increase the probabilities of all classes in the source domain model except for the class C t with the highest probability; the specific implementation method is to minimize the objective function: where is the probability of the source domain model for class C t , the probability of the class C s with the highest probability output by the source domain model is L s represents the loss function on the source domain, the probability of the class C t with the highest probability output by the target domain model is the probability of the class with the second highest probability is B represents the size of each batch, max represents the maximum value function, exp represents the exponential function, and i represents the current batch number in the training process; Step S24: To ensure that the source domain model and the target domain model are not overfitted and to enhance the convergence stability, a regularization term is added to the objective function of the source domain model during training. The regularization term is defined as: where represents the regularization term of the objective function on the source domain; a regularization term is added to the objective function of the target domain model. The regularization term is defined as: where represents the regularization term of the objective function on the target domain; Step S3 specifically includes the following steps: Step S31: After the training batch is greater than b batch , until the training ends, fine-tune the target domain model, where b batch is the batch number set artificially. The specific operation is as follows: Copy a single sample from the target domain 10 times, and perform random preprocessing on each copy according to the preprocessing method in Step S13 to obtain 10 preprocessed samples; Input these 10 preprocessed samples into the target domain model respectively to obtain their outputs. If the number of outputs with the same result is greater than b sc , where b sc is the threshold set artificially, then minimize the cross-entropy loss with this same result as the label; If the number of outputs with the same result is less than b sc , then maximize the cross-entropy loss with the most same result as the label; Step S32: Since during the training process of the source domain model, the mixed samples and labels are used as inputs, and the source domain data features are not extracted sufficiently, the original samples and labels in the source domain are input into the source domain model to fine-tune the source domain model, and the training process is carried out by minimizing the cross-entropy loss; Step S4 specifically includes the following steps: Step S41: When performing inference, input the pictures to be classified into the source domain model and the target domain model respectively. The probability output by the source domain model on class k is The probability output by the target domain model on class k is where 0 < k < class + 1, class is the total number of categories, and calculate for all classes p k The class k corresponding to the maximum value 1 is the inference result; Step S42: If the maximum value of p in all classes k is greater than the threshold Y, then keep class k 1 as the inference result; if the maximum value of p in all classes k is less than the threshold Y, and the class k k with the second-largest value of p in all classes 2 satisfies where is the probability output by the target domain model for class k 2 , is the probability output by the target domain model for class k 1 , then change the inference result to k 2 .
2. The unsupervised domain adaptation method based on bidirectional matching according to claim 1, characterized in that, Step S1 specifically includes the following steps: Step S11: Obtain the publicly available domain adaptation dataset, the office-31 dataset, and obtain the annotations of the training data; Step S12: Use two fixed mixing ratio coefficients λ sd , λ td to perform data mixing; Given a pair of input sample images and a label corresponding to them in the source and target domains: Mix the sample images and labels through the following equation: where λ sd represents the mixing ratio coefficient of the sample image from the source domain, λ td represents the mixing ratio coefficient of the sample image from the target domain, and λ sd + λ td = 1, represents the sample image from the source domain, represents the label of the sample image from the source domain, represents the sample image from the target domain, represents the estimated value of the label of the sample image from the target domain, represents the sample image after mixing, represents the label after mixing, where s represents from the source domain, t represents from the target domain, and i represents which batch process is currently in the training process; Step S13: Perform random transformation on the obtained sample images for data augmentation, and the specific operations are as follows: randomly select 0 to 3 items from (1), (2), and (3) for operations. (1) Randomly rotate the image around the center point, and randomly select the rotation angle from 0° to 360°, and the obtained sample image is cropped to the original image size. (2) Randomly select one of horizontal flipping and vertical flipping for the sample image to perform. (3) Randomly scale the sample image around the center point, and the scaling factor is randomly selected from 0.5 to 1.2, and the obtained sample image is cropped to the original image size.
Citation Information
Patent Citations
Hidden false data injection attack detection method based on deep belief network and transfer learning
CN112560079A
Unsupervised content-preserved domain adaptation method for multiple CT lung texture recognition
US20210390686A1