A cross-domain image classification method and system based on adaptive optimal transmission

Through the adaptive optimal transmission method, deep convolutional networks and fully connected layer networks are used to optimize the solution of image features and classifiers, which solves the problems of insufficient robustness and generalization of cross-domain image classification and realizes efficient image classification in different environments.

CN118334410BActive Publication Date: 2025-10-17SOUTH CHINA UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410342453.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2025-10-17
Estimated Expiration
2044-03-25

AI Technical Summary

Technical Problem

Existing cross-domain image classification technology based on classical optimal transfer theory lacks robustness and generalization in scenarios such as small samples, outliers, noise, and data that does not follow the independent and identically distributed assumption, making it difficult to meet the accuracy and robustness requirements of practical applications.

Method used

An adaptive optimal transmission method is adopted to construct an image feature extractor through a deep convolutional network and an image classifier through a fully connected layer network. Combined with the adaptive optimal transmission distance optimization solution and gradient feedback mechanism, the robustness and generalization of cross-domain image classification are improved.

Benefits of technology

It effectively improves the accuracy and robustness of cross-domain image classification, and can quickly migrate between different environments while maintaining good classification performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118334410B_ABST
    Figure CN118334410B_ABST
Patent Text Reader

Abstract

The application discloses a cross-domain image classification method and system based on adaptive optimal transmission, which comprises the following steps: obtaining original images collected by a source domain and a target domain, and performing image preprocessing and image enhancement; an image feature extractor extracts feature embedding of the image after data enhancement; an image classifier classifies the image according to the feature embedding to obtain a predicted classification label of the image; an adaptive optimal transmission distance between the source domain image set and the target domain image set is solved as a difference degree between the source domain image set and the target domain image set; an image classification loss of the image classifier on the source domain image is calculated; a target function is constructed and iteratively trained to obtain a trained image feature extractor and image classifier; the feature embedding of the image is extracted and the target domain image is classified respectively, and a classification label is output. The application can effectively improve the robustness and generalization of cross-domain image classification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cross-domain image classification, and particularly relates to a cross-domain image classification method and system based on adaptive optimal transport. BACKGROUND

[0002] The cross-domain image classification technology solves the problem of how to quickly migrate an image classification system from an existing environment to a new environment. With the rapid growth of image-related information, image classification is becoming increasingly important in many application fields. Traditional image classification technology requires a large number of labeled image samples, and also requires that the sample distribution of the source domain and the target domain meet the independent and identically distributed assumption, so as to have better results. However, in actual applications, there are sufficient unlabeled image data and a small amount of labeled image data in many fields, and the cost of labeling a large number of image samples is too huge, and in many cases it is even unfeasible. The image data sources of different domains are different, so there is always some difference in the feature distribution or feature space between the domains. For example, the image acquisition equipment and the acquisition conditions have differences, and the images taken in different scenes, different lightings, different angles, etc. are different, and the differences in resolution, expression, action, etc. will also cause the feature distribution to change. The goal of cross-domain image classification is to quickly migrate an image classification system trained in an existing environment (also referred to as a source domain) through a large number of labeled image information to a new environment (also referred to as a target domain).

[0003] At present, cross-domain image classification based on optimal transport theory is one of the most promising research directions in this field. Optimal transport theory studies the difference between two probability distributions (for example: evaluating the difference between the source domain image set and the target domain image set), which is one of the core problems that need to be solved in cross-domain image classification. However, the classic optimal transport problem adopts a probability mass preserving constraint, which severely restricts the performance of the cross-domain image classification system. In the application scenarios of cross-domain image classification, small sample, outlier, noise, long tail effect, data not following the independent and identically distributed assumption, etc. are common. In these scenarios, the mass preserving constraint will distort the transport mapping, severely damaging the robustness and generalization of the cross-domain image classification system. Especially in deep learning, the small batch sampling training paradigm will exacerbate the severity of the problem, and the existing cross-domain image classification technology based on the classic optimal transport theory cannot meet the requirements of many image classification application fields in terms of accuracy and robustness of cross-domain image classification. SUMMARY

[0004] In order to overcome the defects and deficiencies existing in the prior art, the application provides a cross-domain image classification method and system based on adaptive optimal transmission, which can quickly migrate images from an old environment (referred to as a source domain) to a new environment (referred to as a target domain), effectively improve the robustness and generalization of cross-domain image classification, and meet the requirements of cross-domain image classification in terms of accuracy and robustness.

[0005] In order to achieve the above purpose, the application adopts the following technical solutions:

[0006] 1. A cross-domain image classification method based on adaptive optimal transmission, comprising the following steps:

[0007] Obtaining original images collected in a source domain and a target domain, wherein the source domain is provided with a classification label, and the target domain is not provided with a classification label;

[0008] Performing image preprocessing on the original images collected in the source domain and the target domain;

[0009] Performing data enhancement on the preprocessed images;

[0010] Constructing an image feature extractor based on a deep convolutional network, wherein the image feature extractor extracts features from the data-enhanced images to obtain feature embeddings of the images;

[0011] Constructing an image classifier based on a fully connected layer network, wherein the image classifier classifies images according to the feature embeddings of the images to obtain predicted classification labels of the images;

[0012] Based on the feature embeddings of the images and the predicted classification labels, an adaptive optimal transmission distance problem between the image set of the source domain and the image set of the target domain is optimized and solved to obtain an adaptive optimal transmission distance between the image set of the source domain and the image set of the target domain as a difference degree between the image set of the source domain and the image set of the target domain;

[0013] Based on the predicted classification labels of the images, an image classification loss of the image classifier on the images in the source domain is calculated;

[0014] Based on the image classification loss and the difference degree between the image set of the source domain and the image set of the target domain, a target function is constructed, and the image feature extractor and the image classifier are fed back based on the target function to update network parameters and iteratively train to obtain trained image feature extractors and image classifiers;

[0015] Obtaining target domain images, performing feature extraction on the target domain images through the trained image feature extractors to obtain feature embeddings of the images, and classifying the feature embeddings of the target domain images based on the trained image classifiers to output classification labels of the images.

[0016] As a preferred technical solution, the image feature extractor is constructed based on a deep convolutional network, and specifically, a residual deep network ResNet50 is used to construct the image feature extractor.

[0017] As a preferred technical solution, the image classifier is constructed based on a fully connected layer network, and specifically, a single-layer or multi-layer fully connected layer network and a soft-max layer are used to construct the image classifier.

[0018] As a preferred technical solution, the adaptive optimal transport distance between the source domain image set and the target domain image set is solved by optimization, and the adaptive optimal transport distance between the source domain image set and the target domain image set is obtained, which is specifically represented as:

[0019]

[0020] C ij =||g(x i )-g(z j )||-αy i tanh(f(g(z j ))

[0021] Γ ≤ (μ,v)={π∈P(X×Z)|π1 m ≤μ,π T 1 n ≤v}

[0022]

[0023]

[0024] wherein, X represents the source domain image set, Z represents the target domain image set, and W(X, Z) represents the adaptive optimal transport distance between the source domain image set and the target domain image set. ij Γ represents a transport mapping space, π represents a transport mapping, π i represents a transport probability mass from the source domain image x j to the target domain image z ij , C i represents the cost of transporting the unit probability mass from the source domain image x j to the target domain image z i , μ represents a probability measure of the source domain image set, v represents a probability measure of the target domain image set, δx represents a Dirac function of the source domain image, δz represents a Dirac function of the target domain image, p i represents the probability mass of the measure μ at the source domain image x jdenotes the probability mass of the measure v at the target domain image z j denotes the probability measure space, 1 m denotes an m-dimensional unit vector, 1 n denotes an n-dimensional unit vector, g(x i ) denotes the mapping of source domain images to the feature space by the image feature extractor, g(z i ) denotes the mapping of target domain images to the feature space by the image feature extractor, f() denotes the image classifier, alpha is a non-negative coefficient, and tanh denotes the tanh function.

[0025] As a preferred technical solution, the image classification loss of the image classifier on the source domain image is calculated, and is specifically represented as:

[0026]

[0027] Wherein, g() denotes the image feature extractor, f() denotes the image classifier, x i denotes the source domain image, y i denotes the classification label.

[0028] The application also provides a cross-domain image classification system based on adaptive optimal transport, comprising: an original image acquisition module, an image preprocessing module, an image enhancement module, an image feature extractor, an image classifier, an adaptive optimal transport distance calculation module, an image classification loss calculation module, a target function construction module, a training module, and a classification result output module.

[0029] The original image acquisition module is used to acquire original images collected in the source domain and the target domain, the source domain is provided with a classification label, and the target domain is not provided with a classification label.

[0030] The image preprocessing module is used to perform image preprocessing on the original images collected in the source domain and the target domain.

[0031] The image enhancement module is used to perform data enhancement on the preprocessed images.

[0032] The image feature extractor is constructed based on a deep convolutional network, is used for feature extraction on the data-enhanced images, and obtains feature embedding of the images.

[0033] The image classifier is constructed based on a fully connected layer network, is used for image classification according to the feature embedding of the images, and obtains a predicted classification label of the images.

[0034] The adaptive optimal transport distance calculation module is configured to solve an adaptive optimal transport distance problem between the source domain image set and the target domain image set based on the feature embedding of the image and the predicted classification label, and obtain the adaptive optimal transport distance between the source domain image set and the target domain image set as the difference degree between the source domain image set and the target domain image set.

[0035] The image classification loss calculation module is configured to calculate an image classification loss of the image classifier on the source domain image based on the predicted classification label of the image.

[0036] The target function construction module is configured to construct a target function based on the image classification loss and the difference degree between the source domain image set and the target domain image set.

[0037] The training module is configured to perform gradient feedback on the image feature extractor and the image classifier based on the target function, update network parameters, and iteratively train to obtain the trained image feature extractor and image classifier.

[0038] The classification result output module is configured to obtain a target domain image, perform feature extraction on the target domain image by using the trained image feature extractor to obtain a feature embedding of the image, perform classification on the feature embedding of the target domain image by using the trained image classifier, and output a classification label of the image.

[0039] As a preferred technical solution, the image feature extractor is constructed based on a deep convolutional network, and specifically, a residual deep network ResNet50 is used to construct the image feature extractor.

[0040] As a preferred technical solution, the image classifier is constructed based on a fully connected layer network, and specifically, a single-layer or multi-layer fully connected layer network and a soft-max layer are used to construct the image classifier.

[0041] As a preferred technical solution, the adaptive optimal transport distance problem between the source domain image set and the target domain image set is solved to obtain the adaptive optimal transport distance between the source domain image set and the target domain image set, and specifically, the adaptive optimal transport distance between the source domain image set and the target domain image set is represented as:

[0042]

[0043] C ij =||g(x i )-g(z j )||-ay i tanh(f(g(z j )))

[0044] Γ ≤ (μ,v)={π∈P(X×Z)|π1 m ≤μ,π f 1n ≤v}

[0045]

[0046]

[0047] wherein, denotes a source domain image set, denotes a target domain image set, W(X,Z) denotes an adaptive optimal transport distance between the source domain image set and the target domain image set, Γ denotes a transport mapping space, π denotes a transport mapping, π ij denotes a transport probability mass from a source domain image x i at a location to a target domain image z j at a location, C ij denotes a cost of transporting a unit probability mass from a source domain image x i at a location to a target domain image z j at a location, μ denotes a probability measure of the source domain image set, v denotes a probability measure of the target domain image set, denotes a Dirac function of a source domain image, denotes a Dirac function of a target domain image, p i denotes a probability mass of the measure μ at a source domain image x i at a location, q j denotes a probability mass of the measure v at a target domain image z j at a location, P denotes a probability measure space, 1 m denotes an m-dimensional unit vector, 1 n denotes an n-dimensional unit vector, g(x i ) denotes mapping of a source domain image to a feature space by an image feature extractor, g(z i ) denotes mapping of a target domain image to a feature space by an image feature extractor, f() denotes an image classifier, α is a non-negative coefficient, tanh denotes a tanh function.

[0048] As a preferred technical scheme, the image classification loss calculation module is configured to calculate an image classification loss of the image classifier on the source domain image based on a predicted classification label of the image, and the image classification loss is specifically represented as:

[0049]

[0050] wherein, g() denotes an image feature extractor, f() denotes an image classifier, x i denotes a source domain image, y i denotes a classification label.

[0051] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0052] The application adopts adaptive optimal transmission to measure the difference between the source domain image set and the target domain image set, and realizes cross-domain image classification based on adaptive optimal transmission, which can overcome the limitations of the classic optimal transmission theory, effectively improve the robustness and generalization of cross-domain image classification, and meet the requirements of cross-domain image classification in accuracy and robustness. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 A flowchart of the cross-domain image classification method based on adaptive optimal transmission of the application. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical scheme and advantages of the application more clear, the application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and do not limit the application.

[0055] Example 1

[0056] As shown in the figure, the present embodiment provides a cross-domain image classification method based on adaptive optimal transmission, comprising the following steps: Figure 1

[0057] S1: image acquisition; acquiring original images collected from old environment (referred to as source domain) and new environment (referred to as target domain), the source domain image has a classification label (for example, if the image contains pedestrians, the classification label is 'yes', otherwise 'no'), while the target domain image has no classification label;

[0058] In this embodiment, the source domain takes the sunny automatic driving scene as an example, including a large number of manually labeled images, and the target domain takes the snowy automatic driving scene as an example, including images lacking annotation or not manually labeled;

[0059] S2: image preprocessing; the original image is preprocessed by image preprocessing techniques such as cropping and scaling to remove noise information and make the images collected from the old and new environments consistent in dimension;

[0060] In this embodiment, the original image is the original image collected by automatic driving in sunny and snowy scenes, i.e. the automatic driving image;

[0061] S3: data enhancement of image: in order to solve the scarcity of image samples and improve the generalization performance of image classification system, the preprocessed image is also subjected to data enhancement, including but not limited to random slicing, horizontal or vertical flipping, changing lighting conditions, etc.

[0062] ​S4: Feature extraction of the image: a deep convolutional network is used as an image feature extractor to extract features from the data-augmented image, and semantic features are extracted to obtain the feature embedding of the image;

[0063] In this embodiment, the image feature extractor is implemented by using a deep convolutional network, including but not limited to a residual deep network ResNet50, etc. The following takes the residual deep network ResNet50 as an example for introduction: the ResNet50 network has 5 convolutional structures and one average pooling layer. Taking an input color image of 224x224 as an example, first, a convolutional layer conv1 with a number of 64, a convolution kernel size of 7x7, and a step of 2 is used, the layer outputs an image with a size of 112x112 and a channel number of 64; then a 3x3 maximum downsampling pooling layer is used, the layer outputs an image with a size of 56x56 and a channel number of 64; then 4 residual network blocks are stacked, at this time, the output image has a size of 7x7 and a channel number of 2048; finally, an average pooling layer is used to obtain the feature embedding of the image.

[0064] S5: Cross-domain image classification: based on the feature embedding of the image, a fully connected layer network is used as an image classifier to classify the image, and a predicted classification label of the image is obtained (taking the pedestrian recognition task as an example, if the image contains a pedestrian, the classification label is 'yes', otherwise, it is 'no');

[0065] In this embodiment, the image classifier is implemented by using a single-layer or multi-layer fully connected layer network and a soft-max layer;

[0066] S6: Calculate the difference degree between the source domain image set and the target domain image set: based on the feature embedding of the image and the predicted classification label, an adaptive optimal transport model is used to calculate the difference degree between the source domain image set and the target domain image set, which specifically includes:

[0067] The difference degree problem between the source domain image set and the target domain image set is modeled as an adaptive optimal transport problem between image sets, and the specific process is as follows:

[0068] Let the source domain image set be subject to the probability measure the target domain image set be subject to the probability measure where δ is the Dirac function, the measure μ has a probability mass p i at the image x i , the measure v has a probability mass q j at the image z j , p i and q j belong to the probability simplex, that is, and The source domain image set has a classification label, and the classification label set is While the image of the target domain has no classification label. C is a transmission cost function, represents the cost of transmitting unit probability mass from the source domain image x i to the target domain image z j . The transmission mapping is represented by the joint probability π, where π ij represents the transmission probability mass from the source domain image x i to the target domain image z j . P represents the probability measure space, 1 m represents an m-dimensional unit vector, 1 n represents an n-dimensional unit vector. The difference between the source domain image set and the target domain image set is converted into a problem of solving the adaptive optimal transport distance between the image sets, which is specifically represented as:

[0069]

[0070] Where the transmission mapping space is:

[0071] Γ ≤ (μ, v) = {π ∈ P (X × Z) | π m ≤ μ, π T 1 n ≤ v}

[0072] After optimizing the adaptive optimal transport problem between the image sets, the adaptive optimal transport distance W (X, Z) between the image sets is obtained, which is used as the difference between the source domain image set and the target domain image set.

[0073] Specifically, this embodiment takes the automatic driving sunny scene and snow scene image sets as examples, and uses the adaptive optimal transport model to calculate the difference between the automatic driving sunny scene and snow scene image sets.

[0074] Let the automatic driving sunny scene image set obeys the probability measure The automatic driving snow scene image set obeys the probability measure Where δ is the Dirac function, the measure μ has the probability mass p i at the image x i , the measure v has the probability mass g j at the image z j , p i and q j satisfy and The automatic driving sunny scene image set has a classification label, and the classification label set is And the autonomous driving snow scene image set has no classification label. C is the transmission cost function, represents the transmission of unit probability mass from a sunny scene image x i to a snow scene image z j The transmission mapping is represented by the joint probability π, where π ij represents the transmission probability mass from a sunny scene image x i to a snow scene image z j The P represents the probability measure space, 1 m represents the m-dimensional unit vector, 1 n represents the n-dimensional unit vector. The deep network contains an image feature extractor g(x) and a classifier f(x), the image feature extractor maps the image to the feature space, and the image classifier maps the image from the feature space to the classification label space. The following formula is used to calculate the cost of transmitting unit probability mass from a sunny scene image x i to a snow scene image z j :

[0075] C ij =||g(x i )-g(z j )||-αy i tanh(f(g(z j )).

[0076] The construction of the cost function C is to align the ultrasound images in both the feature space and the label space, and it is not comprehensive to consider only the feature or the label space, and the difference between the autonomous driving sunny scene and snow scene image sets is modeled as a problem of solving the adaptive optimal transport distance between the sunny scene and snow scene image sets:

[0077]

[0078] Where α is a non-negative coefficient, and the transmission mapping space is:

[0079] Γ ≤ (μ, v) = {π ∈ P(X × Z) | π1 m ≤ μ, π T 1 n ≤ v}

[0080] After optimizing the adaptive optimal transport problem described above, the adaptive optimal transport distance between the image sets is obtained, which is used as the difference between the autonomous driving sunny scene and snow scene image sets;

[0081] S7: Calculate the image classification loss: Based on the predicted classification label of the image, calculate the classification loss of the image classifier on the source domain image. In this embodiment, the cross entropy loss function is used. As described in step S6, the source domain image set The corresponding classification label set is g() represents the image feature extractor, f() represents the image classifier, and the cross entropy loss function is specifically expressed as:

[0082]

[0083] In this embodiment, the cross entropy loss function is used to calculate the classification loss of the image classifier in the sunny scene of autonomous driving. Of course, the present invention is not limited to using only the cross entropy loss function for calculation, and other classification loss functions are also applicable.

[0084] S8: Calculating the objective function of the neural network: The objective function of the neural network includes the image classification loss and the difference between the source domain image set and the target domain image set. Specifically, the objective function is the sum of the image classification loss and the image set difference, wherein the image classification loss is calculated according to step S7, and the difference between the source domain image set and the target domain image set is calculated according to step S6;

[0085] S9: Neural network gradient feedback: Perform gradient feedback on the above-mentioned deep neural network (including the image feature extractor and the image classifier) ​​according to the objective function of the neural network (obtained in step S8) to update the parameters of the deep neural network;

[0086] S10: Perform neural network training: repeat the above steps S4 to S9 until the neural network converges. For example, if the number of update iterations reaches the maximum number of iterations, it is determined that the neural network has converged.

[0087] S11: Output image classification results: Input the target domain image, perform feature extraction through the above-trained image feature extractor to obtain the feature embedding of the image, and then use the above-trained image classifier to classify the feature embedding of the target domain image and output the classification label of the image.

[0088] In this embodiment, the classification results of the autonomous driving snowy scene images are output, and the autonomous driving snowy scene images are classified using a trained image classifier to output the classification results, such as pedestrian recognition or object classification results.

[0089] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A cross-domain image classification method based on adaptive optimal transmission, characterized in that: The steps include: Obtaining original images collected from a source domain and a target domain, wherein the source domain is provided with a classification label, and the target domain is not provided with a classification label; Perform image preprocessing on the original images collected from the source domain and the target domain; Perform data augmentation on the preprocessed images; An image feature extractor is constructed based on a deep convolutional network. The image feature extractor extracts features from the data-augmented image and obtains feature embedding of the image. An image classifier is constructed based on a fully connected layer network. The image classifier classifies the image based on the feature embedding of the image and obtains the predicted classification label of the image. Based on the feature embedding and predicted classification labels of the image, the adaptive optimal transmission distance problem between the source domain image set and the target domain image set is optimized and solved. The adaptive optimal transmission distance between the source domain image set and the target domain image set is obtained as the difference between the source domain image set and the target domain image set, which is specifically expressed as: C ij =||g(x i )-g(z j )||-αy i tanh(f(g(z j ))) in, represents the source domain image set, represents the target domain image set, W(X, Z) represents the adaptive optimal transmission distance between the source domain image set and the target domain image set, Γ represents the transmission mapping space, π represents the transmission mapping, π ij Represents the image x from the source domain i to the target domain image z j The transmission probability mass at C ij Represents the image x from the source domain i The unit probability mass is transferred to the target domain image z j The cost at , μ represents the probability measure of the source domain image set, ν represents the probability measure of the target domain image set, g(x i ) represents the image feature extractor mapping the source domain image to the feature space, g(z j ) represents the image feature extractor mapping the target domain image to the feature space, f() represents the image classifier, α is a non-negative coefficient, tanh represents the tanh function, y i Indicates the classification label; Based on the predicted classification label of the image, calculate the image classification loss of the image classifier on the source domain image; An objective function is constructed based on the image classification loss and the difference between the source domain image set and the target domain image set. Gradient feedback is performed on the image feature extractor and image classifier based on the objective function to update the network parameters. The trained image feature extractor and image classifier are obtained through iterative training. Obtain the target domain image, extract features through the trained image feature extractor, obtain the feature embedding of the image, classify the feature embedding of the target domain image based on the trained image classifier, and output the classification label of the image.

2. The cross-domain image classification method based on adaptive optimal transmission according to claim 1, characterized in that: The image feature extractor is constructed based on a deep convolutional network, and specifically the residual deep network ResNet50 is used to construct the image feature extractor.

3. The cross-domain image classification method based on adaptive optimal transmission according to claim 1, characterized in that An image classifier is constructed based on a fully connected layer network, specifically using a single-layer or multi-layer fully connected layer network and a soft maximization layer to construct the image classifier.

4. The cross-domain image classification method based on adaptive optimal transmission according to claim 1, characterized in that The transmission mapping space is specifically expressed as: C ≤ (μ,v)={π∈P(X×Z)|π1 m ≤μ,π T 1 n ≤v} in, represents the Dirac function of the source domain image, The Dirac function representing the target domain image, p i Denotes the measure μ in the source domain image x i The probability mass at q j Denotes the measure ν in the target domain image z j The probability mass at , P represents the probability measure space, 1 m represents an m-dimensional unit vector, 1 n Represents an n-dimensional unit vector.

5. The cross-domain image classification method based on adaptive optimal transmission according to claim 1, characterized in that Calculate the image classification loss of the image classifier on the source domain image, specifically expressed as: Among them, g() represents the image feature extractor, f() represents the image classifier, and x i represents the source domain image, y i Represents a classification label.

6. A cross-domain image classification system based on adaptive optimal transmission, characterized in that: include: Original image acquisition module, image preprocessing module, image enhancement module, image feature extractor, image classifier, adaptive optimal transmission distance calculation module, image classification loss calculation module, objective function construction module, training module, classification result output module; The original image acquisition module is used to acquire original images collected from a source domain and a target domain, wherein the source domain is provided with a classification label, and the target domain is not provided with a classification label; The image preprocessing module is used to perform image preprocessing on the original images collected from the source domain and the target domain; The image enhancement module is used to perform data enhancement on the preprocessed image; The image feature extractor is constructed based on a deep convolutional network and is used to extract features from the data-enhanced image to obtain feature embedding of the image; The image classifier is constructed based on a fully connected layer network and is used to classify the image according to the feature embedding of the image to obtain a predicted classification label of the image; The adaptive optimal transmission distance calculation module is used to optimize and solve the adaptive optimal transmission distance problem between the source domain image set and the target domain image set based on the feature embedding and predicted classification label of the image, and obtain the adaptive optimal transmission distance between the source domain image set and the target domain image set as the difference between the source domain image set and the target domain image set, which is specifically expressed as: C ij =||g(x i )-g(z j )||-αy i tanh(f(g(z j ))) in, represents the source domain image set, represents the target domain image set, W(X, Z) represents the adaptive optimal transmission distance between the source domain image set and the target domain image set, Γ represents the transmission mapping space, π represents the transmission mapping, π ij Represents the image x from the source domain i to the target domain image z j The transmission probability mass at C ij Represents the image x from the source domain i Transfer the unit probability mass to the target domain image z j The cost at , μ represents the probability measure of the source domain image set, ν represents the probability measure of the target domain image set, g(x i ) represents the image feature extractor mapping the source domain image to the feature space, g(z j ) represents the image feature extractor mapping the target domain image to the feature space, f() represents the image classifier, α is a non-negative coefficient, tanh represents the tanh function, y i Indicates the classification label; The image classification loss calculation module is used to calculate the image classification loss of the image classifier on the source domain image based on the predicted classification label of the image; The objective function construction module is used to construct an objective function based on the image classification loss and the difference between the source domain image set and the target domain image set; The training module is used to perform gradient feedback on the image feature extractor and the image classifier based on the objective function, update network parameters, and iteratively train to obtain trained image feature extractors and image classifiers; The classification result output module is used to obtain the target domain image, perform feature extraction through the trained image feature extractor to obtain the feature embedding of the image, classify the feature embedding of the target domain image based on the trained image classifier, and output the classification label of the image.

7. The cross-domain image classification system based on adaptive optimal transmission according to claim 6, characterized in that: The image feature extractor is constructed based on a deep convolutional network, specifically using the residual deep network ResNet50 to construct the image feature extractor.

8. The cross-domain image classification system based on adaptive optimal transmission according to claim 6, characterized in that: The image classifier is constructed based on a fully connected layer network, specifically using a single-layer or multi-layer fully connected layer network and a soft maximization layer to construct the image classifier.

9. The cross-domain image classification system based on adaptive optimal transmission according to claim 6, characterized in that: The transmission mapping space is specifically expressed as: C ≤ (μ,ν)={π∈P(X×Z)|π1 m ≤μ,π T 1 n ≤n} in, represents the Dirac function of the source domain image, The Dirac function representing the target domain image, p i Denotes the measure μ in the source domain image x i The probability mass at q j Denotes the measure ν in the target domain image z j The probability mass at , P represents the probability measure space, 1 m represents an m-dimensional unit vector, 1 n Represents an n-dimensional unit vector.

10. The cross-domain image classification system based on adaptive optimal transmission according to claim 6, characterized in that: The image classification loss calculation module is used to calculate the image classification loss of the image classifier on the source domain image based on the predicted classification label of the image. It is specifically expressed as: Among them, g() represents the image feature extractor, f() represents the image classifier, and x i represents the source domain image, y i Represents a classification label.

Citation Information

Patent Citations

  • Unsupervised field adaptive remote sensing image segmentation method based on optimal transmission

    CN113409351A

  • Cross-domain tongue image classification method based on deep hierarchical optimal transmission

    CN116310545A