Model training method and device, image processing method and device, electronic equipment and storage medium

Through mixed enhancement and multiple training methods, the image processing model is trained using the source domain and target domain data and pseudo-label data, which solves the problem of insufficient generalization ability in cross-domain applications and achieves efficient classification in the target domain.

CN120298737APending Publication Date: 2025-07-11HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410033304.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing image processing models have poor generalization capabilities when applied across domains, especially when there is less target domain data, and the classification effect is poor and a large amount of target domain data is required to be retrained.

Method used

By obtaining the source domain sample data of known tags and the target domain sample data of unknown tags, the first classification model is used to generate pseudo-label data, and mixed enhancement is performed. The model is trained by combining the sample data of the source domain and target domain and the tag data, and supervised, unsupervised and adversarial training methods are used to improve the domain adaptability of the model.

Benefits of technology

With a small amount of target domain data, the cross-domain generalization ability and domain adaptability of the model are significantly improved, and the classification accuracy in the target domain is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298737A_ABST
    Figure CN120298737A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device, an image processing method and device, electronic equipment and a storage medium, and the model training method comprises the steps: obtaining source domain sample data of a known label and target domain sample data of an unknown label; obtaining pseudo label data of the target domain sample data through a first classification model based on the target domain sample data; respectively performing hybrid enhancement on the source domain sample data, the target domain sample data, the label data of the source domain sample data and the pseudo label data of the target domain sample data to obtain hybrid sample data and hybrid label data; and training a first classification model based on the source domain sample data, the target domain sample data and the pseudo label data thereof, the mixed sample data and the mixed label data. According to the invention, the domain adaptive capability of the classification model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to a model training, image processing method, device, electronic device, and storage medium. Background Art

[0002] With the vigorous development of neural network and deep learning technologies, image processing technologies have also become increasingly advanced. For example, image data input can be automatically analyzed through an image processing model, including but not limited to classifying elements contained in the image data and generating corresponding labels, etc.

[0003] In the prior art, a large amount of training data is usually required to train the above-mentioned image processing model, and the trained model has poor cross-domain generalization ability and only has a good classification effect on source domain data in the same domain as the training data. If the image processing model is applied to target domain data outside the source domain, a large amount of target domain data is required to retrain the model again. When the target domain data is scarce, a classification model with a good classification effect cannot be trained.

[0004] It should be noted that the above statements are only used to provide background technical information related to this application and do not necessarily constitute prior art. Summary of the Invention

[0005] This application proposes a model training, image processing method, device, electronic device, and storage medium, which can improve the domain adaptation ability of the classification model.

[0006] The first aspect embodiment of this application proposes a model training method, including:

[0007] Obtain source domain sample data with known labels and target domain sample data with unknown labels;

[0008] Based on the target domain sample data, obtain pseudo-label data of the target domain sample data through a first classification model;

[0009] Mix and enhance the source domain sample data, the target domain sample data, the label data of the source domain sample data, and the pseudo-label data of the target domain sample data respectively to obtain mixed sample data and mixed label data;

[0010] Train the first classification model based on the source domain sample data, the target domain sample data and its pseudo-label data, and the mixed sample data and the mixed label data.

[0011] In some embodiments of the present application, the steps of respectively performing hybrid augmentation on the source domain sample data, the target domain sample data, the label data of the source domain sample data, and the pseudo-label data of the target domain sample data to obtain hybrid sample data and hybrid label data include:

[0012] Linearly mix the source domain sample data and the target domain sample data respectively to obtain hybrid sample data;

[0013] Linearly mix the label data of the source domain sample data and the pseudo-label data of the target domain sample data to obtain the hybrid label data corresponding to the hybrid sample data.

[0014] In some embodiments of the present application, the steps of training the first classification model based on the source domain sample data, the target domain sample data and its pseudo-label data, and the hybrid sample data and the hybrid label data include:

[0015] Input each source domain sample in the source domain sample data into the first classification model respectively, and output the source domain classification prediction results corresponding to each source domain sample;

[0016] Input each target domain sample in the target domain sample data into the first classification model respectively, and output the target domain classification prediction results corresponding to each target domain sample;

[0017] Input each hybrid sample in the hybrid sample data into the first classification model respectively, and output the hybrid classification prediction results corresponding to each hybrid sample.

[0018] In some embodiments of the present application, the steps of training the first classification model based on the source domain sample data, the target domain sample data and its pseudo-label data, and the hybrid sample data and the hybrid label data further include:

[0019] Calculate a first loss value based on the source domain classification prediction results respectively corresponding to each source domain sample and the label data of the source domain sample;

[0020] Calculate a second loss value based on the target domain classification prediction results and pseudo-label data respectively corresponding to each target domain sample;

[0021] Calculate a third loss value based on the hybrid classification prediction results and the hybrid label data respectively corresponding to each hybrid sample;

[0022] Adjust the model parameters of the first classification model based on the first loss value, the second loss value, and the third loss value; and loop and execute the step of obtaining the first classification prediction result based on the adjusted model parameters until the first classification model is trained and completed when the first preset convergence condition is reached.

[0023] In some embodiments of the present application, the method further includes:

[0024] Input each source domain sample in the source domain sample data into the second classification model respectively, and output the first domain classification prediction result corresponding to each source domain sample; the second classification model is used to classify the domain to which the input data belongs, including a domain classifier and the backbone network of the first classification model.

[0025] Input each target domain sample in the target domain sample data into the second classification model respectively, and output the second domain classification prediction result corresponding to each target domain sample.

[0026] Based on the first domain classification prediction result and the second domain classification prediction result, perform adversarial training on the backbone network and the domain classifier.

[0027] In some embodiments of the present application, the performing adversarial training on the backbone network and the domain classifier based on the first domain classification prediction result and the second domain classification prediction result includes:

[0028] Calculate a fourth loss value based on the first domain classification prediction result and the second domain classification prediction result.

[0029] Adjust the model parameters of the second classification model based on the fourth loss value.

[0030] Based on the adjusted model parameters, loop and execute the step of obtaining the first domain classification prediction result until the second preset convergence condition is reached.

[0031] In some embodiments of the present application, adjusting the model parameters of the first classification model based on the first loss value, the second loss value, and the third loss value includes:

[0032] Adjust the model parameters of the first classification model based on the first loss value, the second loss value, the third loss value, and the fourth loss value.

[0033] In some embodiments of the present application, the obtaining the pseudo-label data of the target domain sample data through the first classification model based on the target domain sample data includes:

[0034] Perform multiple augmentations on the target domain sample data to obtain the augmented target domain sample data.

[0035] Based on the enhanced target domain sample data, pseudo-label data of the target domain sample data is obtained through the first classification model.

[0036] An embodiment of the second aspect of the present application provides an image processing method, including:

[0037] Obtain target image data to be processed;

[0038] Input the target image data into a pre-trained first classification model, and output the classification result of the target image data;

[0039] Wherein, the first classification model is obtained by using the model training method described in the first aspect.

[0040] An embodiment of the third aspect of the present application provides a model training device, including:

[0041] A sample data acquisition module, configured to acquire source domain sample data with known labels and target domain sample data with unknown labels;

[0042] A pseudo-label data acquisition module, configured to obtain pseudo-label data of the target domain sample data through a first classification model based on the target domain sample data; the first classification model is used to classify each region of the input data;

[0043] A data mixing module, configured to respectively perform mixed enhancement on the source domain sample data and the target domain sample data, as well as the label data of the source domain sample data and the pseudo-label data of the target domain sample data, to obtain mixed sample data and mixed label data;

[0044] A model training module, configured to train the first classification model based on the source domain sample data, the target domain sample data and its pseudo-label data, as well as the mixed sample data and the mixed label data.

[0045] An embodiment of the fourth aspect of the present application provides an image processing device, including:

[0046] A data acquisition module, configured to acquire target image data to be processed;

[0047] An image processing module, configured to input the target image data into a pre-trained classification model and output the classification results of each region of the target image data;

[0048] Wherein, the classification model is the first classification model obtained by using the model training method described in the first aspect.

[0049] An embodiment of the fifth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor runs the computer program to implement the method described in the above first aspect or second aspect.

[0050] An embodiment of the sixth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. The program is executed by a processor to implement the method described in the above first aspect or second aspect.

[0051] The technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0052] In the embodiments of the present application,

[0053] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.

[0055] In the drawings:

[0056] Figure 1 The structural schematic diagram of the first classification model in some embodiments of the present application is shown;

[0057] Figure 2 The flowchart of the model training method provided in some embodiments of the present application is shown;

[0058] Figure 3 The specific flowchart of step S2 in some embodiments of the present application is shown;

[0059] Figure 4 The schematic diagram of the training principle of the first classification model in some embodiments of the present application is shown;

[0060] Figure 5 The specific flowchart of step S3 in some embodiments of the present application is shown;

[0061] Figure 6 The schematic diagram of another training principle of the first classification model in some embodiments of the present application is shown;

[0062] Figure 7 The specific flowchart of step S4 in some embodiments of the present application is shown;

[0063] Figure 8 Shows a further specific process schematic diagram of step S4 in some embodiments of the present application;

[0064] Figure 9 Shows an even more specific process schematic diagram of step S4 in some embodiments of the present application;

[0065] Figure 10 Shows a specific process schematic diagram of step S430 in some embodiments of the present application;

[0066] Figure 11 Shows a structural schematic diagram of a model training device provided in some embodiments of the present application;

[0067] Figure 12 Shows a process schematic diagram of an image processing method provided in some embodiments of the present application;

[0068] Figure 13 Shows a structural schematic diagram of an image processing device provided in some embodiments of the present application;

[0069] Figure 14 Shows a schematic diagram of an electronic device provided in an embodiment of the present application;

[0070] Figure 15 Shows a schematic diagram of a storage medium provided in an embodiment of the present application. Detailed implementation manners

[0071] Hereinafter, the exemplary embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be completely conveyed to those skilled in the art.

[0072] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meanings understood by those skilled in the art to which the present application belongs.

[0073] In the field of image processing, an unknown-label image can be classified by a trained classification model to achieve tasks such as land parcel classification, face recognition, and vehicle statistics of images. However, this classification model usually requires a large amount of training data for training, and the trained model has poor cross-domain generalization ability and only has good classification effects on source domain data in the same domain as the training data. If this image processing model is applied to target domain data outside the source domain, a large amount of target domain data is required to retrain the model. When the target domain data is scarce, the trained classification model often does not have good classification effects on the target domain data.

[0074] Source domain data can be understood as the data used to train the classification model, which can be a large amount of easily obtainable data with labels; target domain data can be understood as the data to which the classification model is to be applied, which can be a small amount of data that is not easily obtainable and has no labels or very few labels. There are often certain differences between source domain data and target domain data, such as changes in data distribution and lack of features. For example, satellite remote sensing image data is used as source domain data, and unmanned aerial vehicle (UAV) remote sensing image data is used as target domain data. The differences between the two include, but are not limited to: UAV remote sensing image data can collect data at fixed times and locations, can capture higher-quality images, and can obtain real-time remote sensing images; while satellite data can only collect relevant data based on satellite flight plans, cannot collect data at fixed times and locations, cannot obtain real-time remote sensing images, and has lower image capture quality. The differences between source domain data and target domain data will have a negative impact on the performance of the classification model, resulting in poor classification effects of the classification model on target domain data.

[0075] To solve the above problems, this embodiment proposes a model training method. After obtaining source domain sample data and target domain sample data, the pseudo-label data of the target domain sample data can be obtained through a first classification model first. Then, the source domain sample data and the target domain sample data, as well as the label data of the source domain sample data and the pseudo-label data of the target domain sample data, are respectively subjected to mixed augmentation to obtain mixed sample data and mixed label data. Based on the source domain sample data, the target domain sample data and its pseudo-label data, as well as the mixed sample data and the mixed label data, the first classification model is trained. In this way, the first classification model is supervised-trained through the source domain sample data to improve the classification accuracy of the first classification model on the source domain data; the first classification model is unsupervised-trained through the target sample data and its corresponding pseudo-label data to improve the domain adaptation ability of the first classification model; and by respectively performing mixed augmentation on the source domain sample data and the target domain sample data, as well as their label data, and then training the first classification model with the mixed sample data and the mixed label data, the cross-domain generalization ability of the first classification model can be effectively improved, so that a classification model with stronger cross-domain generalization ability and better domain adaptation ability can be trained with a small amount of target domain sample data.

[0076] In this embodiment, the first classification model may include a backbone network and a main classifier. Among them, the backbone network realizes feature embedding of the input data based on an appropriate spatial mapping relationship, maps the high-dimensional input data, including but not limited to the source domain sample data and the target domain sample data in this embodiment, to a low-dimensional space, and obtains a multi-dimensional real-valued vector. The main classifier can be used to classify each element in the input data and generate corresponding labels. Denote the main network as f and the main classifier as g, then the first classification model can be expressed as h = fog, and the specific structure is as Figure 1 shown. And this embodiment does not make specific limitations on the specific structures of the backbone network and the main classifier, as long as the backbone network can realize the above-mentioned feature embedding function and the main classifier can realize the above-mentioned classification function. For example, including but not limited to convolutional deep neural networks or sequence models (Transformer) based on attention mechanisms, such as Residual Network (ResNet), Vision Transformer (ViT) network. The main classifier can be any generative classifier, such as a Naive Bayes model, a Gaussian mixture model, and a Hidden Markov model, etc.

[0077] The execution subject of this model training method can be any computer device, cluster device, or cloud processing platform that loads the first classification model and can perform classification processing on image data. This embodiment does not make specific limitations on this.

[0078] The following elaborates in detail on the model training method provided by the embodiments of the present application with reference to the accompanying drawings.

[0079] Please refer to Figure 2 , which is a schematic flowchart of the model training method provided by the embodiment of the present application. As Figure 2 shown, the model training method may include the following steps:

[0080] Step S1, obtain source domain sample data with known labels and target domain sample data with unknown labels.

[0081] Among them, the source domain sample data with known labels and its label data are regarded as a whole set S, and the set S can be denoted as Among them, represents the elements in the set S, and each element is a set of training data. is the source domain sample, is the source domain sample corresponding label data. One source domain sample can correspond to a set of label data, and m s is the total number of source domain sample data. The target domain sample data with unknown labels is regarded as a set T, and the set T can be denoted as Among them, represents the elements of the set T, and each element is a training sample, and m t is the total number of target domain sample data.

[0082] Specifically, the source domain sample data with known labels can be obtained from existing databases, and can include but are not limited to open-source satellite remote sensing data with known labels. The target domain sample data with unknown labels can be collected by an image acquisition device according to requirements, and can include but are not limited to drone remote sensing data with unknown labels.

[0083] Step S2, based on the target domain sample data, through the first classification model, obtain the pseudo-label data of the target domain sample data.

[0084] Among them, the pseudo-label data can be the label data generated by directly classifying the target domain sample data using an untrained first classification model. The accuracy of this label data is poor, so it is called pseudo-label data.

[0085] In some embodiments, as Figure 3 shown, step S2 may include the following specific steps: Step S21, perform multiple augmentations on the target domain sample data to obtain the augmented target domain sample data; Step S22, based on the augmented target domain sample data, through the first classification model, obtain the pseudo-label data of the target domain sample data.

[0086] In this embodiment, the total number of target domain sample data can be much less than that of source domain sample data. To further improve the classification effect of the first classification model in the target domain, the target domain sample data can be enhanced to obtain more target domain sample data. Then, the enhanced target domain sample data is input into the untrained first classification model to obtain corresponding pseudo-label data. In this way, the original sample data in the target domain can be enhanced and expanded by the enhanced target domain sample data and the corresponding pseudo-label data, so as to further improve the robustness and domain adaptation ability of the first classification model.

[0087] Specifically, as Figure 4 shown, any training sample in the target domain can be randomly data-augmented K times through the first augmentation module. The data augmentation methods can include but are not limited to cropping, rotation, color shift, etc. The augmented sample data can be denoted as where x represents an augmented sample after random augmentation of any training sample, K is a natural number greater than or equal to 1, representing the total number of augmentations of the training sample i,k and also the number of augmented samples based on the training sample

[0088] The original sample data and the augmented sample data in the target domain are subjected to feature mapping through the backbone network to obtain corresponding feature embedding vectors. The set of feature embedding vectors corresponding to all sample data in the target domain can be denoted as This backbone network can include a normalization function, and through the normalization function, such as Softmax, the above-mentioned set of feature embedding vectors can be transformed into a scalar representing the probability distribution. For example, the above-mentioned set of feature embedding vectors can be transformed into probability values

[0089] Based on the self-consistency requirement, the probability distributions of the sample data after different random data augmentations of the same original sample data should be as consistent as possible in the feature space. Therefore, the pseudo-label corresponding to the augmented sample data is defined as:

[0090] Using the above-mentioned randomly augmented sample data and the corresponding pseudo-label data to train the first classification model, due to the randomness and self-consistency of the augmented samples, the sensitivity of the first classification model to random noise can be reduced, thereby enhancing the robustness of the first classification model.

[0091] ​​​It should be noted that the target domain sample data in the following steps S3 and S4 may include the original sample data of the target domain and the enhanced sample data. The pseudo-label data of the target domain sample data may also include the pseudo-label data corresponding to the original sample data of the target domain and the pseudo-label data corresponding to the enhanced sample data.

[0092] Step S3: Mix and enhance the source domain sample data and the target domain sample data, as well as the label data of the source domain sample data and the pseudo-label data of the target domain sample data respectively, to obtain mixed sample data and mixed label data.

[0093] Among them, mix and enhance can be understood as mixing the sample data of the source domain and the target domain to obtain new sample data, and mixing the label data and pseudo-labels corresponding to the sample data before mixing in the same way to obtain new label data, and forming new training samples with the new sample data and the corresponding new label data.

[0094] In this embodiment, the sample data of the source domain and the target domain and the corresponding label data (including the label data of the source domain and the pseudo-label data of the target domain) are fused in the same way respectively to obtain new training samples. Using these new training samples to train the first classification model can improve the cross-domain generalization ability of the first classification model.

[0095] In some embodiments, as Figure 5 shown, step S3 may include the following specific steps: Step S31: Linearly mix the source domain sample data and the target domain sample data respectively to obtain mixed sample data; Step S32: Linearly mix the label data of the source domain sample data and the pseudo-label data of the target domain sample data to obtain the mixed label data corresponding to the mixed sample data.

[0096] Among them, linear mixing means that there is a certain linear relationship between the data after mixing and the data before mixing.

[0097] In this embodiment, using the linear mixing strategy, the source domain sample data and the target domain sample data, as well as the label data of the source domain sample data and the pseudo-label data of the target domain sample data, are linearly mixed respectively to obtain mixed sample data and the corresponding mixed label data. In this way, the constraint relationship between the mixed sample data and the sample data before mixing is the same as the constraint relationship between the corresponding mixed label data and the label data before mixing, which can further improve the cross-domain generalization ability of the first classification model.

[0098] Specifically, as mentioned above, the source domain data including the source domain sample and its corresponding label can be expressed as The target domain data including the target domain sample and its corresponding pseudo-label can be expressed as AsFigure 6 As shown in the following formula, any source domain sample can be mixed with any target domain sample through the second enhancement module based on the linear mixing strategy to obtain a mixed sample. Similarly, based on the following linear mixing strategy, the label data corresponding to any source domain sample and the pseudo-label data corresponding to any target domain sample are linearly mixed to obtain the mixed label corresponding to the mixed sample.

[0099]

[0100]

[0101] where λ is a linear parameter, which is greater than 0 and less than 1 and satisfies a symmetric beta distribution.

[0102] It can be understood that the above linear mixing method is only one implementation method for data mixing enhancement in this embodiment. This embodiment is not limited thereto. As long as the sample data of the source domain and the target domain and the corresponding label data can be mixed respectively to achieve data enhancement and improve the cross-domain generalization ability of the classification model.

[0103] Step S4: Train the first classification model based on the source domain sample data, the target domain sample data and their pseudo-label data, as well as the mixed sample data and the mixed label data.

[0104] In practical applications, each training sample in the source domain sample data, each training sample in the target domain sample data, and each training sample in the mixed sample data can be respectively input into the backbone network of the first classification model, and then the main classifier classifies the output result of the backbone network. Then, the loss value is calculated based on the classification result, and the first classification model is adjusted based on the loss value.

[0105] In some embodiments, as Figure 7 shown, step S4 may include the following specific steps: Step S411: Input each source domain sample in the source domain sample data into the first classification model, and output the source domain classification prediction result corresponding to each source domain sample; Step S412: Input each target domain sample in the target domain sample data into the first classification model, and output the target domain classification prediction result corresponding to each target domain sample; Step S413: Input each mixed sample in the mixed sample data into the first classification model, and output the mixed classification prediction result corresponding to each mixed sample.

[0106] ​​Among them, the source domain samples can be understood as the image data of the source domain without labeled data; the target domain samples can be understood as the image data of the target domain; the mixed samples can be understood as the image data obtained after performing function operations on the image data of the source domain and the image data of the target domain.

[0107] In this embodiment, inputting the source domain samples into the first classification model can perform supervised training on the first classification model based on the labeled data of the source domain samples. Inputting the target domain samples into the first classification model can perform unsupervised training based on the pseudo-labels of the target domain samples. Inputting the mixed samples into the first classification model can perform weak supervised training based on the mixed label data of the mixed sample data. Thus, multi-dimensional training can be performed on the first classification model, improving the classification accuracy of the first classification model in the target domain, that is, further improving the cross-domain generalization ability of the first classification model.

[0108] It can be understood that the above steps S411 to S413 are only for indicating three different steps and do not constitute a limitation on the time sequence of the three steps. That is, in this embodiment, it is not limited to the input order of the source domain samples, the target domain samples, and the mixed samples, nor is it necessary to wait for the classification prediction result of the previous input sample to be output before inputting the next sample, as long as the corresponding classification prediction result can be obtained.

[0109] Specifically, as Figure 8 shown, step S4 may further include the following specific steps: step S421, calculating a first loss value based on the source domain classification prediction results respectively corresponding to each source domain sample and the label data of the source domain samples; step S422, calculating a second loss value based on the target domain classification prediction results and pseudo-label data respectively corresponding to each target domain sample; step S423, calculating a third loss value based on the mixed classification prediction results and mixed label data respectively corresponding to each mixed sample; step S440, adjusting the model parameters of the first classification model based on the first loss value, the second loss value, and the third loss value; and based on the adjusted model parameters, returning to the step of obtaining the first classification prediction result and looping until the first preset convergence condition is reached to obtain the trained first classification model.

[0110] Among them, the first loss is denoted as L cls , and it can be the cross-entropy loss, that is The second loss is denoted as L consisitency , which is a self-consistent regularization loss constructed according to the target domain samples and the corresponding pseudo-labels and is also the cross-entropy loss, that is The third loss is denoted as L domain , which is a cross-domain generalization loss obtained based on the mixed samples and their mixed label data and is also the cross-entropy loss, that is The first preset convergence condition may, but is not limited to, the loss value tending to 0. The preset convergence conditions for the three loss values may be the same or different, and this embodiment does not make specific limitations in this regard.

[0111] In this embodiment, based on the source domain classification prediction result, the target domain classification prediction result, and the mixed classification prediction result respectively, the corresponding loss values are calculated, and then based on the loss values, the first classification model is adjusted. After the model is adjusted, the new training samples are input into the first classification model to obtain new loss values, and then based on the new loss values, the first classification type after the previous adjustment is adjusted again for parameter adjustment. This is executed in a loop until the latest loss values all reach their respective first preset convergence conditions, so as to obtain a first classification model with better classification effect in the target domain.

[0112] In some embodiments, as Figure 9 shown, step S4 may further include the following steps: step S414, input each source domain sample in the source domain sample data into the second classification model respectively, and output the first domain classification prediction result corresponding to each source domain sample; the second classification model is used to classify the domain to which the input data belongs, including the domain classifier and the backbone network of the first classification model; step S424, input each target domain sample in the target domain sample data into the second classification model respectively, and output the second domain classification prediction result corresponding to each target domain sample; step S430, based on the first domain classification prediction result and the second domain classification prediction result, perform adversarial training on the backbone network and the domain classifier.

[0113] Among them, the second classification model may share the backbone network with the first classification model, or two identical backbone networks may be set, and this embodiment does not make specific limitations in this regard.

[0114] In this embodiment, the domain classifier is also used to perform adversarial training on the backbone network and the domain classifier. The accuracy of the domain classifier can be verified based on the classification result of the domain classifier for the source domain samples, and then based on the classification result of the domain classifier for the target domain samples, the accuracy of the spatial mapping relationship of the backbone network is trained. If the accuracy of the classification result of the domain classifier for the target domain samples is higher, it means that the spatial mapping accuracy of the backbone network is lower. Performing adversarial training in this way can further improve the accuracy of the backbone network.

[0115] Specifically, as Figure 10 shown, step S430 may include the following steps: step S431, calculate the fourth loss value based on the first domain classification prediction result and the second domain classification prediction result; step S432, adjust the model parameters of the second classification model based on the fourth loss value; step S433, based on the adjusted model parameters, return to the step of obtaining the first domain classification prediction result and execute in a loop until the second preset convergence condition is reached.

[0116] Among them, the fourth loss is denoted as L align , which is the distribution alignment loss. It can perform cross-domain representation consistency constraints on source domain samples and target domain samples. It can be, but is not limited to, KL divergence constraints. Then the fourth loss The setting of the second preset convergence condition is similar to that of the first preset convergence condition. It can be, but is not limited to, that the loss value tends to 0.

[0117] In this embodiment, based on the first domain classification prediction result and the second domain classification prediction result, the corresponding fourth loss value is calculated. Then, based on this loss value, the second classification model (including the backbone network of the domain classifier and the first classification model) can be adjusted. Then, the new training samples are input into the second classification model to obtain a new fourth loss value. Then, based on this new fourth loss value, the second classification type adjusted last time is adjusted again for parameter adjustment. This is executed in a loop until the latest fourth loss value reaches the second preset convergence condition, so as to obtain a backbone network with better spatial mapping.

[0118] Furthermore, as Figure 9 shown, the model parameters of the first classification model can also be adjusted based on the fourth loss value by performing step S440'. In this way, based on the four loss values, the final loss can be obtained as Loss = L consisitency + L domain - L align + L cls . This can not only adjust the first classification model from multiple angles, but also perform adversarial training on the backbone network based on the results of adversarial training, so as to obtain a first classification model with self-consistent regularization, strong inter-domain generalization ability, and good cross-domain representation consistency.

[0119] It can be understood that this embodiment does not limit that adversarial training must be performed separately and the fourth loss value must be calculated. The source domain training samples, target domain training samples, and mixed training samples can also be input into the backbone network in sequence (without distinguishing the order). Then, part of the output results of the backbone network of the source domain samples are input into the main classifier, and part are input into the domain classifier; similarly, part of the output results of the backbone network of the target domain samples are input into the main classifier, and the other part are input into the domain classifier; the output results of the backbone network of the mitigation samples are all input into the main classifier, so as to obtain the first loss value, the second loss value, the third loss value, and the fourth loss value at one time. Then, the model parameters of the first classification model are adjusted based on the first loss value, the second loss value, the third loss value, and the fourth loss value.

[0120] In summary, the model training method provided in this embodiment can, after obtaining the source domain sample data and the target domain sample data, first obtain the pseudo-label data of the target domain sample data through the first classification model. Then, the source domain sample data and the target domain sample data, as well as the label data of the source domain sample data and the pseudo-label data of the target domain sample data, are respectively subjected to mixup augmentation to obtain the mixed sample data and the mixed label data. Furthermore, based on the source domain sample data, the target domain sample data and its pseudo-label data, as well as the mixed sample data and the mixed label data, the first classification model is trained. In this way, the first classification model is supervised-trained with the source domain sample data to improve the classification accuracy of the first classification model on the source domain data; the first classification model is unsupervised-trained with the target sample data and its corresponding pseudo-label data to improve the domain adaptation ability of the first classification model; and by respectively performing mixup augmentation on the source domain sample data and the target domain sample data, as well as their label data, and then training the first classification model with the mixed sample data and the mixed label data, the cross-domain generalization ability of the first classification model can be effectively improved. Thus, in the case of a small amount of target domain sample data, a classification model with stronger cross-domain generalization ability and better domain adaptation ability can be trained.

[0121] Some embodiments of the present application also provide a model training device for implementing the above model training method. As Figure 11 shown, the model training device may include:

[0122] A sample data acquisition module for acquiring a plurality of source domain sample data with known labels and a plurality of target domain sample data with unknown labels;

[0123] A pseudo-label data acquisition module for obtaining the pseudo-label data of the target domain sample data based on the target domain sample data through the first classification model; the first classification model is used to classify each region of the input data;

[0124] A data mixing module for respectively performing mixup augmentation on the source domain sample data and the target domain sample data, as well as the label data of the source domain sample data and the pseudo-label data of the target domain sample data, to obtain the mixed sample data and the mixed label data;

[0125] A model training module for training the first classification model based on the source domain sample data, the target domain sample data and its pseudo-label data, as well as the mixed sample data and the mixed label data.

[0126] It can be understood that the model training device provided in this embodiment and the model training method provided in the embodiments of the present application are based on the same inventive concept, and can at least achieve the same beneficial effects as the model training method. Moreover, various implementation manners of the model training method embodiments are also equally applicable to the embodiments of this model training device, and will not be elaborated herein.

[0127] Some embodiments of the present application also provide an image processing method. As Figure 12 shown, the method may include the following steps:

[0128] Step S100: Obtain target image data to be processed.

[0129] Step S200: Input the target image data into a pre-trained first classification model, and output the classification result of the target image data.

[0130] Among them, the first classification model is obtained by using the above-mentioned model training method. The target image data may be the image data of the above target domain. Outputting the classification result of the target image data means generating the label data of the target image data.

[0131] It can be understood that the image processing method provided in this embodiment uses the first classification model obtained by the above-mentioned model training method for image processing, and can at least achieve the same functions and effects as the first classification model. Moreover, various implementation manners of the model training method embodiment are also applicable to the embodiment of this image processing method, and will not be elaborated here.

[0132] Some embodiments of the present application also provide an image processing device for implementing the above-mentioned image processing method. As Figure 13 shown, the device may include:

[0133] A data acquisition module for obtaining target image data to be processed;

[0134] An image processing module for inputting the target image data into a pre-trained classification model and outputting the classification results of each region of the target image data;

[0135] Among them, the first classification model is obtained by using the above-mentioned model training method. The target image data may be the image data of the above target domain. Outputting the classification result of the target image data means generating the label data of the target image data.

[0136] It can be understood that the image processing device provided in this embodiment uses the first classification model obtained by the above-mentioned model training method for image processing, and can at least achieve the same functions and effects as the first classification model. Moreover, various implementation manners of the model training method embodiment are also applicable to the embodiment of this image processing method, and will not be elaborated here.

[0137] It should be noted that the data involved in this application (including but not limited to data for model training, stored data, displayed data, etc.) are all information and data that have been authorized by users or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0138] The embodiments of this application also provide an electronic device to execute the above-mentioned model training method or image processing method. Please refer to Figure 14 , which shows a schematic diagram of an electronic device provided by some embodiments of this application. As Figure 14 shown, the electronic device 4 includes: a processor 400, a memory 401, a bus 402, and a communication interface 403. The processor 400, the communication interface 403, and the memory 401 are connected through the bus 402; a computer program that can run on the processor 400 is stored in the memory 401, and when the processor 400 runs the computer program, it executes the model training method or image processing method provided by any of the foregoing embodiments of this application.

[0139] Among them, the memory 401 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 403 (which can be wired or wireless), a communication connection between this device network element and at least one other network element can be realized, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0140] The bus 402 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 401 is used to store programs. After receiving an execution instruction, the processor 400 executes the program, and the model training method or image processing method disclosed in any of the foregoing embodiments of this application can be applied to the processor 400 or implemented by the processor 400.

[0141] The processor 400 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method may be completed by the integrated logic circuit of the hardware in the processor 400 or the instructions in the form of software. The above-mentioned processor 400 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 401, and the processor 400 reads the information in the memory 401 and combines its hardware to complete the steps of the above method.

[0142] The electronic device provided by the embodiments of the present application and the model training method or image processing method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by them.

[0143] The embodiments of the present application also provide a computer-readable storage medium corresponding to the model training method or image processing method provided in the foregoing embodiments. Please refer to Figure 15 which shows that the computer-readable storage medium is an optical disc 50, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the model training method or image processing method provided in any of the foregoing embodiments.

[0144] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.

[0145] The embodiments of the present application also provide a computer program product, including a computer program, which is executed by a processor to implement the model training method or the image processing method in any of the above embodiments.

[0146] The computer-readable storage medium and the computer program product provided in the above embodiments of the present application are all based on the same inventive concept as the model training method or the image processing method provided in the embodiments of the present application, and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0147] It should be noted that:

[0148] In the specification provided herein, a large number of specific details are set forth. However, it is understood that the embodiments of the present application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0149] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the following schematic: that the claimed present application requires more features than are expressly recited in each of the claims. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.

[0150] In addition, those skilled in the art will appreciate that, although some embodiments described herein include certain features included in other embodiments but not others, the combination of features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0151] The above is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A model training method, characterized in that, The method includes: Obtaining source domain sample data with known labels and target domain sample data with unknown labels; Based on the target domain sample data, obtaining pseudo-label data of the target domain sample data through a first classification model; Performing hybrid augmentation on the source domain sample data, the target domain sample data, the label data of the source domain sample data, and the pseudo-label data of the target domain sample data respectively to obtain hybrid sample data and hybrid label data; Training the first classification model based on the source domain sample data, the target domain sample data and its pseudo-label data, and the hybrid sample data and the hybrid label data.

2. The method according to claim 1, wherein The performing hybrid augmentation on the source domain sample data, the target domain sample data, the label data of the source domain sample data, and the pseudo-label data of the target domain sample data respectively to obtain hybrid sample data and hybrid label data includes: Performing linear mixing on the source domain sample data and the target domain sample data respectively to obtain hybrid sample data; Performing linear mixing on the label data of the source domain sample data and the pseudo-label data of the target domain sample data to obtain hybrid label data corresponding to the hybrid sample data.

3. The method according to claim 1, wherein The training the first classification model based on the source domain sample data, the target domain sample data and its pseudo-label data, and the hybrid sample data and the hybrid label data includes: Inputting each source domain sample in the source domain sample data into the first classification model respectively, and outputting source domain classification prediction results corresponding to each source domain sample; Inputting each target domain sample in the target domain sample data into the first classification model respectively, and outputting target domain classification prediction results corresponding to each target domain sample; Inputting each hybrid sample in the hybrid sample data into the first classification model respectively, and outputting hybrid classification prediction results corresponding to each hybrid sample.

4. The method according to claim 3, wherein The training the first classification model based on the source domain sample data, the target domain sample data and its pseudo-label data, and the hybrid sample data and the hybrid label data further includes: Calculating a first loss value based on the source domain classification prediction results respectively corresponding to each source domain sample and the label data of the source domain sample; Calculating a second loss value based on the target domain classification prediction results and pseudo-label data respectively corresponding to each target domain sample; Calculating a third loss value based on the hybrid classification prediction results and the hybrid label data respectively corresponding to each hybrid sample; Adjusting the model parameters of the first classification model based on the first loss value, the second loss value and the third loss value; and based on the adjusted model parameters, returning to the step of obtaining the first classification prediction results and executing the loop until a first preset convergence condition is reached to obtain a trained first classification model.

5. The method according to claim 4, wherein The method further includes: Input each source domain sample in the source domain sample data into the second classification model respectively, and output the first domain classification prediction result corresponding to each source domain sample; the second classification model is used to classify the domain to which the input data belongs, and includes a domain classifier and the backbone network of the first classification model; Input each target domain sample in the target domain sample data into the second classification model respectively, and output the second domain classification prediction result corresponding to each target domain sample; Based on the first domain classification prediction result and the second domain classification prediction result, perform adversarial training on the backbone network and the domain classifier.

6. The method according to claim 5, wherein The performing adversarial training on the backbone network and the domain classifier based on the first domain classification prediction result and the second domain classification prediction result includes: Calculate a fourth loss value based on the first domain classification prediction result and the second domain classification prediction result; Adjust the model parameters of the second classification model based on the fourth loss value; Based on the adjusted model parameters, return to the step of obtaining the first domain classification prediction result and loop until the second preset convergence condition is reached.

7. The method according to claim 6, wherein Adjusting the model parameters of the first classification model based on the first loss value, the second loss value and the third loss value includes: Adjust the model parameters of the first classification model based on the first loss value, the second loss value, the third loss value and the fourth loss value.

8. The method according to claim 1, characterized in that, The obtaining the pseudo-label data of the target domain sample data through the first classification model based on the target domain sample data includes: Perform multiple augmentations on the target domain sample data to obtain the augmented target domain sample data; Based on the augmented target domain sample data, obtain the pseudo-label data of the target domain sample data through the first classification model.

9. An image processing method, characterized in that, Includes: Obtain the target image data to be processed; Input the target image data into a pre-trained first classification model, and output the classification result of the target image data; Wherein, the first classification model is obtained by using the model training method according to any one of claims 1-9.

10. A model training device, characterized in that, Includes: A sample data acquisition module, configured to acquire source domain sample data with known labels and target domain sample data with unknown labels; A pseudo-label data acquisition module, configured to obtain the pseudo-label data of the target domain sample data through a first classification model based on the target domain sample data; the first classification model is used to classify each region of the input data; A data mixing module, configured to perform mixed augmentation on the source domain sample data and the target domain sample data, and the label data of the source domain sample data and the pseudo-label data of the target domain sample data respectively to obtain mixed sample data and mixed label data; A model training module, configured to train the first classification model based on the source domain sample data, the target domain sample data and its pseudo-label data, and the mixed sample data and the mixed label data.

11. An image processing apparatus, characterized in that, Includes: A data acquisition module, configured to acquire the target image data to be processed; An image processing module, configured to input the target image data into a pre-trained classification model and output classification results of each region of the target image data; Wherein, the classification model is a first classification model obtained by using the model training method described in any one of claims 1-8.

12. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the method described in any one of claims 1-9.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method described in any one of claims 1-9.