Domain-adaptive image recognition and classification method in passive continuous scene, storage medium and electronic equipment
By adopting semantic reconstruction and invariance quantization methods in passive continuous domain adaptation scenarios, the catastrophic forgetting and domain offset problems of the model when facing new environments are solved, and higher adaptability performance and accuracy are achieved.
Patent Information
- Application Number
- CN202510055430.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-14
AI Technical Summary
In passive continuous domain adaptation scenarios, when the model faces changing data flows and new environments, it is prone to catastrophic forgetting of the knowledge learned and it is difficult to effectively deal with the domain offset problem.
Semantic reconstruction method is used to obtain a negative sample data set similar to the semantics of the source domain, and the samples are given weights through invariance quantization, and high-quality samples are selected to train the model.
Through semantic reconstruction and invariance quantization methods, the model can reduce catastrophic forgetting in a passive environment, mitigate domain shifts, improve long-term adaptation performance, and significantly improve the average accuracy on Office-31, Office-Home, and DomainNet datasets.
Smart Images

Figure CN120088538A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning, and particularly to a domain adaptation image recognition and classification method, a storage medium, and an electronic device in a source-free continual scenario. Background Art
[0002] Although current deep learning models have shown great potential in various applications, in practical applications, the models are affected by distribution shift. That is, when the data of the training source domain and the target domain of application come from different distributions, when the model is applied to the new target domain, the performance will show a significant decline. Therefore, the unsupervised domain adaptation method of aligning the data distributions of the source domain and the target domain has become an important means to solve this challenge.
[0003] However, most unsupervised domain adaptation methods require access to the source data, which is not applicable in the case of privacy protection or data corruption. Thus, the source-free unsupervised domain adaptation method that directly adapts in the target domain without accessing the source data has also begun to be studied. Nevertheless, when the model is in the process of continual learning, the source-free unsupervised domain adaptation method will suffer from catastrophic forgetting of the learned knowledge. In addition, although the traditional class-incremental learning method can learn new knowledge while retaining the existing knowledge, thus solving the forgetting problem, it cannot effectively cope with the domain shift problem.
[0004] Therefore, the source-free continual adaptation scenario that can continuously learn and adapt to the new environment in the face of changing data streams without accessing the source data is very important for practical applications. For example, the training data of the recognition system in autonomous driving may initially be mainly based on the urban environment. When the vehicle is driving in rural areas or mountainous areas, due to the obvious difference between the target domain and the source domain (urban environment), the model needs to continuously adapt to the new scenario and continuously learn new data in an unsupervised manner during the driving process. At the same time, due to privacy or storage capacity limitations, the source data of the urban environment during training may not be available. This scenario actually requires the model to be able to continuously learn and adapt to the object classification task in the new environment without accessing the source data. Summary of the Invention
[0005] The present invention proposes a domain adaptation image recognition and classification method, a storage medium, and an electronic device in a source-free continual scenario. A negative sample data set similar in semantics to the source domain is obtained by semantic reconstruction (Semantic Restruction). The sample invariance quantization method (Source-Free Continual Adaptation, SFCA) in source-free continual adaptation assigns weights to samples, and high-quality samples are selected to train the model, so as to solve the problem of performance degradation caused by domain shift in existing source-free continual domain adaptation technology.
[0006] In an embodiment of the present invention, a domain adaptation image recognition and classification method in a source-free continual scenario includes the following steps:
[0007] S1. Perform augmentation processing on the source domain data, input the augmented samples into the source domain model to obtain prediction results, and construct a confusion matrix according to the prediction results;
[0008] S2. Perform semantic recombination on the confusion matrix to obtain a negative sample data set similar in semantics to the source domain; use the negative sample data set and the source domain data to train an image quality quantization model, and obtain an image quality quantization model Mn with a divided decision boundary;
[0009] S3. Perform invariance quantization processing on the image samples of the target data set, input any image sample of the target data set into the quality quantization model Mn to obtain the probabilities of the input image sample being assigned to each image category; then quantify the quality of the inherent features of the input image sample in each image category to measure the quality of the input image sample, and assign different weights to the input image sample according to the quality;
[0010] S4. Train the target model. At the end of each training round, the target model extracts features from the target domain image data in the current incremental task, calculates the feature centers of each image category, and calculates the distance from the feature of each target domain image data to the feature center. Select a preset number of images with smaller distances for each image category for storage;
[0011] S5. Accumulate the weights of all image samples classified into each image category to obtain a weight statistic value; compare the weight statistic value with a preset threshold. When the weight statistic value of any image category is greater than the preset threshold, it is determined that the image samples of the corresponding image category exist in a large number in the current batch, so as to identify the image category most likely to be included in each batch and obtain an image sample label;
[0012] S6. Optimize the total loss function of the target model, train and optimize the target model to obtain a final target model for recognizing target domain images in a source-free continual adaptation scenario;
[0013] Among them, the total loss function includes the cross-entropy loss of the target domain images in the current incremental task, the contrastive learning loss of the target domain images of different categories, and the cross-entropy loss of the images stored in step S4.
[0014] Based on the same inventive concept, the domain adaptation image recognition and classification system in the passive continuous scenario in the embodiments of the present invention is implemented based on the above classification method. The classification system includes:
[0015] Data augmentation module: used to perform augmentation processing on the source domain data set and construct a confusion matrix;
[0016] Semantic recombination module: perform semantic recombination on the confusion matrix to obtain a negative sample data set similar to the source domain semantics; use the negative sample data set and the source domain data to train an image quality quantization model to obtain an image quality quantization model Mn with a divided decision boundary;
[0017] Invariance quantization module: perform invariance quantization processing on the image samples of the target data set, input any image sample of the target data set into the quality quantization model Mn to obtain the probabilities of the input image sample being assigned to each image category; then quantify the quality of the inherent features of the input image sample in each image category to measure the quality of the input image sample, and assign different weights to the input image sample according to the quality;
[0018] Memory storage module: train the target model. At the end of each training round, the target model extracts features from the target domain image data in the current incremental task, calculates the feature centers of each image category, and calculates the distance from the feature of each target domain image data to the feature center, and stores a preset number of images with smaller distances for each image category;
[0019] Source label recognition module: accumulate the weights of all image samples classified into each image category to obtain a weight statistical value; compare the weight statistical value with a preset threshold. When the weight statistical value of any image category is greater than the preset threshold, it is determined that the image samples of the corresponding image category exist in a large number in the current batch, so as to identify the most likely image category included in each batch and obtain an image sample label;
[0020] Prediction and recognition module: optimize the total loss function of the target model, train and optimize the target model to obtain a final target model for recognizing target domain images in a passive continuous adaptive scenario.
[0021] The storage medium of the embodiment of the present invention stores computer instructions, and when the computer instructions are executed by a processor, the steps of the above classification method are implemented.
[0022] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory. The computer program can run on the processor. When the processor executes the computer program, it performs the steps of the above classification method.
[0023] Compared with the prior art, the present invention has the following advantages and effects:
[0024] 1. By adopting semantic reconstruction, a negative sample data set similar to the source domain semantics is constructed, enabling the model to be trained on both positive and negative samples simultaneously, thereby accurately configuring the decision boundary; by involving negative samples in the model training, the model's recognition ability for negative classes is enhanced, improving the robustness of the model and enabling it to maintain high adaptation performance in various domain adaptation tasks.
[0025] 2. By adopting Invariance Quantification, based on the probability output by the pre-trained model, the target samples are invariantly quantified, weights are assigned to the samples, enabling the model to select high-quality samples for training, ensuring that both catastrophic forgetting can be reduced and domain shift can be alleviated in a source-free environment; effectively reducing catastrophic forgetting that occurs during the continuous learning process of the model and improving long-term adaptation performance.
[0026] 3. Experiments show that compared with the prior art, the average accuracy of the present invention on the Office-31, Office-Home, and DomainNet data sets has increased by 4.9%, 7.3%, and 10.2% respectively. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a comparison schematic diagram for domain adaptation, source-free domain adaptation, incremental domain adaptation, and source-free incremental domain adaptation.
[0028] Figure 2 It is a flowchart of the domain adaptation image recognition and classification method in the source-free continuous scenario according to an embodiment of the present invention;
[0029] Figure 3 It is a general framework schematic diagram of the domain adaptation image recognition and classification system in the source-free continuous scenario according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The present invention will be further described in detail below with reference to the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0031] Embodiment
[0032] In this embodiment, a domain adaptation image recognition and classification method in a passive continuous scenario is proposed, which is mainly achieved through semantic recombination and sample invariance quantization in passive continuous adaptation. Its purpose is to accurately classify unlabeled images in the target domain into corresponding categories without accessing the source domain dataset, while continuously learning new target domain images and reducing catastrophic forgetting of historical images, thereby improving the robustness and generalization ability of the model.
[0033] Figure 1 The comparison of domain adaptation, source-free domain adaptation, class incremental domain adaptation, and source-free class incremental domain adaptation is shown. Different boxes represent different tasks, and the lock represents that the corresponding arrow data cannot be used. Domain adaptation can use the source domain dataset and only needs to train for one task; class incremental domain adaptation can use the source domain dataset and needs to train for multiple tasks; source-free domain adaptation cannot use the source domain dataset and only needs to train for one task; source-free class incremental domain adaptation cannot use the source domain dataset and needs to train for multiple tasks. Some of the terms involved are explained as follows:
[0034] (a) Domain Adaptation (DA)
[0035] The source domain data contains labeled images, and the source domain model is obtained through training. The target domain data is unlabeled images, and the image categories are the same as those of the source domain images. The source domain model needs to continue training on the target domain images to achieve the goal of accurately recognizing unlabeled target domain images. In the training stage, both unlabeled target domain data and labeled source domain data can be used for training.
[0036] (b) Source-Free Domain Adaptation (SFDA)
[0037] The source domain data contains labeled images, and the source domain model is obtained through training. The target domain data is unlabeled images, and the image categories are the same as those of the source domain images. During the process of the source domain model continuing to train on the target domain images, the source domain data cannot be used anymore, and only the target domain data can be used for training.
[0038] (c) Class Incremental Domain Adaptation (CIDA)
[0039] Similar to (a), the source domain data can still be used in the target domain training stage after the model is trained. For unlabeled target domain images, although the source domain data can be accessed during training, the target domain data can only be used in batches. The training is divided into multiple stages, which are carried out sequentially. Only the partial target domain data of the current stage can be used for training in each stage, and the target domain data of the previous stage cannot be used anymore.
[0040] (d) Source-Free Continual Adaptation (SFCA)
[0041] This scenario is applicable to the present invention. The source domain data includes labeled images, and a source domain model is obtained through training with a model. During the target domain training phase, only the source domain model 3 can be used, and the source domain data cannot be used. For the target domain data, the training is divided into multiple stages, which are carried out sequentially. Only partial target domain data of the current stage can be used for training in each stage, and the target domain data of the previous stage cannot be used anymore.
[0042] The target domain data is unlabeled images. During the training phase, the labeled source domain data cannot be used, and only the source domain model can be used. And the same as in (c), the target domain data can only be used in batches. Finally, the performance of the model on all target domain data needs to be tested.
[0043] As Figure 2 shown, the steps of this embodiment specifically include:
[0044] Step 1: Perform augmentation processing on the source domain data, input the augmented samples into the source domain model to obtain prediction results, and construct a confusion matrix according to the prediction results.
[0045] Source domain data X t After data augmentation, augmented images are generated that are different from their respective categories in local features (operations such as rotating, scaling, and color-changing the original image). By inputting the augmented samples into the source domain model Mp, prediction results are obtained. Due to the similarity in features of these augmented samples (for example, a color-changed pencil image is similar to a pen image), the model will misclassify some of the augmented images. According to the prediction results, a confusion matrix can be constructed. For any image category, the confusion matrix can provide the classification statistical results of all image samples of this image category, so that the category with the most misclassifications (i.e., the category most similar to this class) other than its own category can be selected.
[0046] Step 2: Semantic reconstruction: Perform semantic recombination on the confusion matrix to obtain a negative sample data set similar to the source domain semantics; use the negative sample data set and the source domain data to train an image quality quantization model, and obtain an image quality quantization model Mn with a divided decision boundary.
[0047] Through the confusion matrix in Step 1, after selecting all categories with high similarity to the source domain categories, multiple groups of similar category pairs can be obtained where y s is the source category, is the category most similar to the source category. In a set of similar category pairs, randomly select one image from each pair. The two images will be cropped along the same dividing line to obtain half of the original image. The different halves of the two images are spliced and recombined to obtain a negative example dataset (i.e., a negative sample dataset).
[0048] By jointly using the negative example dataset and the source domain data to train the image quality quantization model, an image quality quantization model Mn with a divided decision boundary can be obtained. Through the output result of the image quality quantization model for the sample and the invariance quantization method, the quality of the sample can be measured during the training process, making the target model more inclined to train high-quality images, thereby improving the model performance.
[0049] Step 3, Invariance Quantization: Perform invariance quantization processing on the image samples of the target dataset. Input any image sample of the target dataset into the quality quantization model Mn to obtain the probabilities of the input image sample being assigned to each image category; then quantify the quality of the inherent features of the input image sample in each image category to measure the quality of the input image sample, and assign different weights to the input image sample according to the quality.
[0050] Invariance quantization is a method for measuring image quality. Input any image sample into the quality quantization model Mn obtained in Step 2, and the probabilities of the input image sample being assigned to each image category can be obtained. When the probability of a certain image category is high, for example, exceeding a preset probability value, it means that the input image sample contains a large number of inherent features (Invariance) unique to that image category. For example, an image recognized as the category "person" has clear facial features, torso, and limbs. By quantifying the quality of the inherent features of the input image sample in each image category through a formula, the quality of the input image sample can be measured.
[0051] Specifically, in this step, different weights are assigned to the image samples according to the positions of the image samples in the feature space. The formula is as follows:
[0052]
[0053]
[0054] where x t is the image sample to be quantified, is the probability that the image sample is assigned to the i-th category, C s is the source domain label set, C n is the negative class (i.e., negative sample) label set, is the sum of the probabilities that the image sample is assigned to the negative class categories, W n (x t) is an indicator to measure how far an image sample is outside the decision boundary (the farther the image sample is from the decision boundary, the larger this indicator is), W p (x t ) is an indicator to measure how close an image sample is to the class feature center (the closer the image sample is to the feature center, the smaller this indicator is). By combining W p (x t ) and W n (x t ) we get W(x t ) to measure the overall quality of the image sample.
[0055] When an image sample is outside the decision boundary, it means that the image sample lacks obvious classification features and the quality of the image sample is poor. The value of W n (x t ) is high, and the value of W(x t ) is small. At this time, a lower weight is assigned to the image sample. On the contrary, the closer the image sample is to the class feature center, the higher the quality of the image sample. The value of W p (x t ) is high, and the value of W(x t ) is large. The image sample gets a higher weight.
[0056] Step 4, Memory Storage: Train the target model. At the end of each training round, the target model extracts features from the target domain image data in the current incremental task, calculates the feature centers of each image class, and calculates the distance from the feature of each target domain image data to the feature center, and selects N images with smaller distances for storage for each image class.
[0057] After the invariance quantization process in Step 3, the target model needs to store the target domain image data in the current incremental task. After each training, the target model first uses the feature extractor G(·) to extract the features of each target domain image data, then averages the features of all target domain image data with the same image class, calculates the feature center μ of each image class, calculates the distance from the feature of each target domain image data to the feature center μ respectively, and selects N images (i.e., images with higher quality) that are closest to the feature center (i.e., with smaller distances) for storage for each image class. The formula is as follows:
[0058]
[0059] where μ k is the feature center of the k-th image class, is the i-th image to be stored for the k-th image class, D k is the image data set of the k-th image class, x i is the image to be selected for storage, x jIs a stored image.
[0060] In the next training task, the target model is trained only on the target domain images and the stored images in the current incremental task to maintain the memory of old knowledge, thereby alleviating the catastrophic forgetting problem faced by the target model in continual learning and improving the generalization ability of the target model.
[0061] Step 5, source label recognition: Accumulate the weights of all image samples classified into each image category to obtain a weight statistical value; compare the weight statistical value with a preset threshold. When the weight statistical value of any image category is greater than the preset threshold, it is determined that the image samples of the corresponding image category exist in large numbers in the current batch, so as to identify the image category most likely to be included in each batch and obtain the image sample label.
[0062] To further reduce the impact of noise in unlabeled images on the performance of the target model, in this embodiment, a weight W(x t ) is obtained for each image sample through the formula in Step 3 and the quality quantization model Mn, and the weights of all image samples classified into each image category are accumulated to obtain the statistical value v k of the weights, where k represents the k-th image category. When the weight statistical value v k of any image category is greater than the preset threshold α, it is considered that the image samples of the k-th category exist in large numbers in the current batch, so that the image category most likely to be included in each batch can be identified. This method can reduce the possibility of misidentifying images, thereby reducing training noise and improving the sample quality.
[0063] Step 6, optimize the total loss function of the target model, train and optimize the target model to obtain the final target model for identifying target domain images in a source-free continuous adaptation scenario.
[0064] Combining the above modules, the total loss function L includes the cross-entropy loss L CE of the target domain images in the current incremental task, the contrastive learning loss L con of target domain images of different categories, and the cross-entropy loss L bank
[0065] L = L CE + λL con + L bank
[0066]
[0067] Among them, X t is the image sample space of the current incremental task, x bank is the sample space of the stored images in Step 4, w(x t) is the corresponding image sample weight, y i is the corresponding image sample label, C b is the stored image sample label set, c + is the feature center of the category to which the image sample belongs, c i is the feature center of the remaining image categories; τ and λ are weight parameters used to control the influence of the contrast learning loss. By training the target model and optimizing the above loss function, the target model can accurately identify target domain images in a source-free continuous adaptation scenario.
[0068] Through the above steps, this embodiment achieves accurate classification of image sample categories on the unlabeled target domain. Compared with the currently best-performing method, the average accuracy rates on the Office-31, Office-Home, and DomainNet datasets are increased by 4.9%, 7.3%, and 10.2% respectively, and it shows higher stability in dealing with domain transfer and catastrophic forgetting problems.
[0069] It can be seen that based on the source-free continuous adaptation (SFCA) framework, this embodiment proposes a new invariance quantization. By selecting pictures of categories with higher similarity for image cropping and reconstruction, it divides the decision boundary for the model, uses the invariance quantization method to screen out pictures of higher quality, stores historical high-quality images to mitigate catastrophic forgetting of historical data, and identifies possible image categories in the unsupervised target domain, achieving accurate classification of unlabeled target domain pictures. This embodiment solves the problem in the prior art that it is impossible to simultaneously alleviate domain transfer and catastrophic forgetting in a source-free continuous adaptation scenario, thereby improving the robustness and generalization ability of the model and achieving accurate recognition of images.
[0070] As Figure 3 shown, this embodiment also proposes a domain adaptation image recognition and classification system in a source-free continuous scenario. The overall framework mainly includes a target model, a memory storage module (Memory Buffer), a source label recognition module (Prototype Identification), a semantic recombination module (Semantic Restruction), and an invariance quantization module (Invariance Quantification).
[0071] Data augmentation module: used to perform augmentation processing on the source domain dataset (Source Datasets) and construct a confusion matrix. The source domain dataset is used to train the source domain model Mp and construct a negative example dataset.
[0072] Semantic Restruction module: Semantically restructure the confusion matrix to obtain a negative sample data set with semantics similar to the source domain; use the negative sample data set and the source domain data to train the image quality quantization model to obtain an image quality quantization model Mn with a divided decision boundary.
[0073] When generating the negative sample data set, select categories with high similarity from the confusion matrix, crop and recombine the images of similar categories to construct the negative sample data set.
[0074] The quality quantization model Mn is jointly trained by the source domain data and the negative example data, and can quantify the image quality through the decision boundary divided during the training process.
[0075] Invariance Quantification module: Perform invariance quantification processing on the image samples of the target data set. Input any image sample of the target data set into the image quality quantization model Mn to obtain the probabilities of the input image sample being assigned to each image category; then quantify the quality of the inherent features of the input image sample in each image category to measure the quality of the input image sample, and assign different weights to the input image sample according to the quality.
[0076] In this embodiment, the Invariance Quantification module assigns different weights to different pictures according to the output of the model Mn (i.e., the position of the image features in space) and the quantization formula. When the image features are outside the decision boundary of the model, it means that the image quality is low, and the formula assigns a lower weight to the image. On the contrary, the weight of a high-quality image is higher. This makes the model more inclined to learn high-quality images and improves the robustness of the model.
[0077] Memory Buffer module: Train the target model. At the end of each training round, the target model extracts features from the target domain image data in the current incremental task, calculates the feature centers of each image category, and calculates the distance from the feature of each target domain image data to the feature center, and selects a preset amount of images with smaller distances for each image category for storage.
[0078] In this embodiment, the Memory Buffer module is used to store images of previous batches and continue training in subsequent tasks to alleviate the problem of catastrophic forgetting of historical images by the model due to the passage of time.
[0079] Prototype Identification Module: Accumulate the weights of all image samples classified into each image category to obtain a weight statistical value; compare the weight statistical value with a preset threshold. When the weight statistical value of any image category is greater than the preset threshold, it is determined that the image samples of the corresponding image category exist in a large number in the current batch, so as to identify the image category most likely to be included in each batch and obtain an image sample label.
[0080] Since the target domain images are unlabeled, the Prototype Identification Module can identify the image categories most likely to exist in each batch, thereby reducing label noise and improving the sample quality.
[0081] Prediction and Identification Module: Optimize the total loss function of the target model, train and optimize the target model to obtain the final target model, so as to identify the target domain images in a source-free continuous adaptation scenario.
[0082] Based on the same inventive concept, this embodiment also provides a storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the steps S1 - S6 of the classification method of this embodiment are implemented.
[0083] Correspondingly, this embodiment also provides an electronic device, including a memory, a processor, and a computer program stored on the memory, where the computer program can run on the processor. When the processor executes the computer program, the steps S1 - S6 of the classification method of this embodiment are executed.
[0084] Essentially, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs, etc., which can store program codes.
[0085] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A domain-adaptive image recognition and classification method in a passive continuous scene, characterized in that: The following steps are involved: S1. Enhance the source domain data, input the enhanced samples into the source domain model to obtain the prediction results, and construct a confusion matrix based on the prediction results; S2, semantically reorganize the confusion matrix to obtain a negative sample dataset with similar semantics to the source domain; The image quality quantization model is trained using the negative sample data set and the source domain data to obtain the image quality quantization model Mn with the decision boundary divided; S3, performing invariance quantization processing on the image samples of the target data set, inputting any image sample of the target data set into the quality quantization model Mn, obtaining the probability of the input image sample being classified into each image category; then quantizing the quality of the inherent features of the input image sample in each image category to measure the quality of the input image sample, and assigning different weights to the input image sample according to the quality; S4, training the target model. At the end of each training round, the target model extracts features from the target domain image data in the current incremental task, calculates the feature center of each image category, and calculates the distance from the feature of each target domain image data to the feature center, and selects a preset amount of images with a smaller distance for each image category for storage; S5, accumulating the weights of all image samples classified into each image category to obtain a weight statistic value; Compare the weighted statistical value with the preset threshold. When the weighted statistical value of any image category is greater than the preset threshold, it is determined that a large number of image samples of the corresponding image category exist in the current batch, so as to identify the image category most likely to be included in each batch and obtain the image sample label; S6. Optimizing the total loss function of the target model, training and optimizing the target model, and obtaining the final target model to recognize the target domain image in a passive continuous adaptive scenario; The total loss function includes the cross entropy loss of the target domain image in the current incremental task, the contrastive learning loss of target domain images of different categories, and the cross entropy loss of the image stored in step S4.
2. The domain adaptation image recognition and classification method according to claim 1, characterized in that: The process of obtaining the negative sample data set in step S2 includes: After selecting all categories with high similarity to the source domain categories through the confusion matrix, multiple groups of similar category pairs are obtained. where y s is the source category, is the category most similar to the source category; randomly select an image from a set of similar category pairs, and the two images will be cropped along the same dividing line to obtain half of the original image. The different halves of the two images will be spliced and recombined to obtain a negative sample data set.
3. The domain adaptation image recognition and classification method according to claim 1, characterized in that: The invariance quantization process of step S3 includes: Input any image sample into the quality quantification model Mn to obtain the probability of the input image sample being classified into each image category; when the probability of a certain image category exceeds the preset probability value, the quality of the inherent features of the input image sample in each image category is quantified through the formula to measure the quality of the input image sample.
4. The domain adaptation image recognition and classification method according to claim 3, characterized in that: The formula is: Among them, x t is the image sample to be quantized, is the probability of the image sample being classified into the i-th image category, C s is the source domain label set, C n is the negative sample label set, is the probability and W of image samples being classified into negative sample categories n (x t ) is an indicator to measure whether the image sample falls outside the decision boundary, W p (x t ) is an indicator to measure the proximity of image samples to the center of the category feature. By combining W p (x t ) and W n (x t ) to get W(x t ) to measure the overall quality of image samples.
5. The domain adaptation image recognition and classification method according to claim 1, characterized in that: The calculation formula for selecting and storing the image in step S4 is: where μ k is the feature center of the k-th image category, is the i-th image to be stored for the k-th image category, D k is the image dataset of the kth image category, x i For the image to be stored, x j For stored images.
6. The domain adaptation image recognition and classification method according to claim 1, characterized in that: The total loss function L in step S6 and the cross entropy loss L of the target domain image in the current incremental task CE , contrastive learning loss L for target domain images of different categories con And the cross entropy loss L of the stored image bank They are: L=L CE +λL con +L bank Among them, X t is the image sample space of the current incremental task, X bank is the sample space of the image stored in step S4, w(x t ) is the corresponding image sample weight, y i is the corresponding image sample label, C b is the stored image sample label set, c + is the feature center of the category to which the image sample belongs, c i are the feature centers of the remaining image categories; τ and λ are weight parameters used to control the influence of contrastive learning loss.
7. A domain-adaptive image recognition and classification system in a passive persistent scene, characterized in that: The classification system is implemented based on any one of the classification methods in claims 1 to 6, and comprises: Data enhancement module: used to enhance the source domain data set and construct a confusion matrix; Semantic Reorganization Module: semantically reorganize the confusion matrix to obtain a negative sample dataset with similar semantics to the source domain; use the negative sample dataset and source domain data to train the image quality quantization model to obtain the image quality quantization model Mn with the decision boundary divided; Invariance quantization module: performs invariance quantization processing on the image samples of the target data set, inputs any image sample of the target data set into the quality quantization model Mn, obtains the probability of the input image sample being classified into each image category; then quantifies the quality of the inherent features of the input image sample in each image category to measure the quality of the input image sample, and assigns different weights to the input image sample according to the quality; Memory storage module: trains the target model. At the end of each training round, the target model extracts features from the target domain image data in the current incremental task, calculates the feature center of each image category, and calculates the distance from the feature of each target domain image data to the feature center, and selects a preset amount of images with a smaller distance for each image category for storage; Source label recognition module: The weights of all image samples classified into each image category are accumulated to obtain a weight statistic value; the weight statistic value is compared with a preset threshold value. When the weight statistic value of any image category is greater than the preset threshold value, it is determined that a large number of image samples of the corresponding image category exist in the current batch, so as to identify the image category most likely to be included in each batch and obtain the image sample label; Prediction and recognition module: optimizes the total loss function of the target model, trains and optimizes the target model, and obtains the final target model to recognize the target domain image in a passive continuous adaptive scenario.
8. A storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the steps of the classification method described in any one of claims 1 to 6 are implemented.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the computer program can be run on the processor, wherein: When the processor executes the computer program, the processor performs the steps of the classification method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Cross-domain remote sensing scene classification and retrieval method based on self-supervised contrast learning
CN115471739A
Image segmentation model training method, image processing method, system and electronic equipment
CN116188481A
Domain generalization method based on text regularization
CN116628555A
Image classification method, system and device based on active domain self-adaption and medium
CN116630708A
Self-adaptive model training method and device and terminal equipment
CN119004236A
Cited By
Real-time detection and tracking method and device for water athletes and medium
CN120747171A
A real-time detection and tracking method, device and medium for water sports players
CN120747171B
Strong enhancement-based instance-class contrast learning unsupervised field adaptive image classification method
CN121482454A