Single-domain generalization method based on texture enhancement and form enhancement

By combining texture enhancement and morphological enhancement in the single-domain generalization method, diverse virtual samples are generated, which solves the shortcomings of existing methods in terms of morphological differences and significantly improves the generalization ability of the classification model.

CN120219178APending Publication Date: 2025-06-27GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510410016.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing single-domain generalization method has achieved certain results in texture differences, but ignores the domain deviation problem caused by morphological differences, which leads to insufficient coverage of sample domains by the generated virtual samples, limiting the generalization performance of the model.

Method used

Using a single domain generalization method based on texture enhancement and morphological enhancement, virtual samples with diverse texture styles and morphological diversity are generated by training the texture enhancement module and the morphological enhancement module, and used in combination to train the classification model.

Benefits of technology

By enriching the sample domain of virtual samples, the model's adaptability to diversified data is improved, the receiving field of model training is expanded, and the generalization ability of classification models is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219178A_ABST
    Figure CN120219178A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a texture enhancement and morphological enhancement-based single-domain generalization method, which comprises the following steps of: acquiring a first source domain image and a second source domain image, and preprocessing the first source domain image and the second source domain image; inputting the first source domain image into a trained texture enhancement module to obtain a corresponding first virtual sample; inputting the first source domain image and the second source domain image into a trained morphological enhancement module to obtain a corresponding second virtual sample; training a classification model by using the first source domain image, the first virtual sample and the second virtual sample, designing a loss function to optimize a training process, and when the number of training times reaches a specified number of times, stopping training to obtain a trained classification model; and inputting a target domain image into the trained classification model, and outputting a classification result. According to the method, the domain difference problem caused by the texture difference and the form difference is considered at the same time, the sample domain of the generated virtual sample is enriched, and therefore the generalization ability of the classification model is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a single-domain generalization method based on texture enhancement and morphological enhancement. Background Art

[0002] Currently, the methods for single-domain generalization generate virtual samples with diverse texture styles by enhancing the texture of source-domain samples, so as to enrich the sample diversity of the model input data, and thus reduce the domain difference from the target-domain data. For example, L2D (Learning to Diversify for Single Domain Generalization) generates samples with invisible styles beyond the original distribution through the style complementarity of different convolutional kernels. PDEN (Progressive Domain Expansion Network for Single Domain Generalization) uses adversarial training to maximize the contrastive learning loss to create a group of samples with domain-invariant features. This method also increases the coverage rate of the sample domain and improves the integrity of the sample domain by gradually expanding the generated sample group. Although these methods have better solved the domain difference problem caused by texture differences, they ignore the domain bias problem caused by morphological differences. Therefore, the virtual samples generated by these methods still have insufficient coverage of the sample domain and can only limitedly improve the generalization performance of the model. Summary of the Invention

[0003] The purpose of the present invention is to overcome the problems existing in the prior art, and provide a single-domain generalization method based on texture enhancement and morphological enhancement. The present invention can enrich the sample domain of the generated virtual samples, and use diverse virtual samples to train a classification model, thereby effectively improving the generalization ability of the classification model.

[0004] To achieve the above purpose, the present invention provides a single-domain generalization method based on texture enhancement and morphological enhancement, and the method includes the following steps: Obtain a first source-domain image and a second source-domain image, and perform preprocessing; Input the preprocessed first source-domain image into a trained texture enhancement module to obtain a corresponding first virtual sample; Input the preprocessed first source-domain image and the preprocessed second source-domain image into a trained morphological enhancement module to obtain a corresponding second virtual sample; Use the first source-domain image, the first virtual sample, and the second virtual sample to train a classification model, and design a loss function to optimize the training process. When the number of training times reaches the specified number, stop training to obtain a trained classification model; Input the target domain image into the trained classification model to output the classification result.

[0005] Further, the preprocessing includes at least one of normalization, rotation, scaling, and translation.

[0006] Further, input the preprocessed first source domain image into the trained texture enhancement module to obtain the corresponding first virtual sample. The texture enhancement module includes: an encoder, an adaptive instance normalization layer, a fully connected layer, and a decoder. Specifically, the encoder is used to obtain texture features based on the input preprocessed first source domain image, and the texture features include at least one of color, illumination, shadow, edge direction, local contrast, and object surface attributes; the fully connected layer is used to map random Gaussian noise to generate the mean and variance of the style image; the adaptive instance normalization layer is used to replace the mean and variance of the texture features with the mean and variance of the style image to obtain the second texture features, and the decoder is used to obtain the first virtual sample based on the second texture features.

[0007] Further, input the preprocessed first source domain image and the preprocessed second source domain image into the morphological enhancement module to obtain the corresponding second virtual sample. The morphological enhancement module includes a multi-scale morphological coordinate structure and a deformator. The multi-scale morphological coordinate structure is used to obtain a set of morphological key point coordinates based on the preprocessed first source domain image and the preprocessed second source domain image; the deformator is used to obtain the second virtual sample based on the set of morphological key point coordinates and the first source domain image.

[0008] Further, the multi-scale morphological coordinate structure is used to obtain a set of morphological key point coordinates based on the preprocessed first source domain image and the preprocessed second source domain image, specifically including: (1). Input the preprocessed first source domain image and the preprocessed second source domain image into an encoder with n feature extraction layers for feature extraction, and denote the feature maps extracted by the nth feature extraction layer of the encoder as 、 , perform similarity comparison on the corresponding pixels of 、 , and denote the pixel coordinates with similarity greater than the threshold as the set of similar coordinates between the two feature maps of the nth feature extraction layer and the set of morphological coordinate and the set of morphological coordinates ; (2). Upsample the set of morphological coordinates 、 of the two feature maps of the nth layer to map and obtain two feature maps corresponding to the (n - 1)th layer 、 Coordinate sets of the same size ; (3) Compare the corresponding pixels of the two feature maps in the (n - 1)-th layer , and record the pixel coordinates with similarity greater than the threshold as the similarity coordinate set between the two feature maps in the (n - 1)-th layer ; (4) Perform dot product weighted aggregation on the similarity coordinate set and the coordinate set to obtain the morphological coordinate set between the two feature maps in the (n - 1)-th layer ; (5) Repeat steps (2), (3), and (4) until the morphological coordinate set of the two feature maps in the first feature extraction layer is obtained . Process the morphological coordinate set in the first layer to obtain the morphological key point coordinate set with the same size as the first source domain image and the second source domain image .

[0009] Further, the similarity comparison is to calculate the square difference between the pixel coordinates at the corresponding positions of the two feature maps, and the calculation formula is as follows:

[0010] where , are the pixel coordinates of the two feature maps , respectively.

[0011] Further, the threshold is determined by the following formula:

[0012] where is the mean value of the feature map , is the variance of the feature map .

[0013] Further, the deformer is used to obtain the second virtual sample according to the morphological key point coordinate set and the first source domain image, and the specific process is as follows: Randomly deform the first source domain image according to the morphological key point coordinate set to obtain the initially deformed image. Specifically, assume that the morphological key point coordinate set has a total of key point coordinates, that is , and randomly generate a displacement vector set with , translate the set of key point coordinates based on the set of displacement vectors to obtain the initially deformed image; Generate a dense flow field according to the pixel coordinates of the first source domain image and the pixel coordinates of the deformed image, and determine parameters based on the dense flow field. The parameters include a weight coefficient, a polynomial constructed to solve a linear equation system, and an augmented vector; Establish an inverse mapping function based on the parameters; Generate the color of each pixel through bilinear sampling of the inverse mapping function and the pixel coordinates of the first source domain image, and output the final deformed image, that is, the second virtual sample.

[0014] Furthermore, the inverse mapping function is determined by the following formula:

[0015] where 、 、 are the weight coefficient, the polynomial constructed to solve the linear equation system, and the augmented vector respectively, represents the position of the pixel in the deformed image, is the radial basis function matrix, specifically .

[0016] Furthermore, the loss functions for training the texture enhancement module, training the morphological enhancement module, and training the classification model are the same, and both include a maximized entropy loss function and a minimized cross-entropy loss function. The maximized entropy loss function is determined by the following formula:

[0017] where 、 、 are the probabilities of predicting the occurrence of the i-th category event in the first source domain image, the first virtual sample, and the second virtual sample respectively, and i is the index of the category event; the minimized cross-entropy loss function is determined by the following formula:

[0018] where 、 、 are the prediction outputs of the first source domain image, the first virtual sample, and the second virtual sample respectively, 、 、 are the category labels of the corresponding images, 、 、 are the predicted categories of the classification model 、 、 probability.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention uses a trained texture enhancement module and a trained morphological enhancement module to obtain a first virtual sample and a second virtual sample respectively, enriching the sample domain of the generated virtual samples. Then, by combining the first virtual sample and the second virtual sample to train a classification model, the data presentation method is enriched from multiple perspectives, making it more adaptable to the changes in various test data sets, expanding the receptive field of the classification model training, and achieving the purpose of improving the model generalization. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flowchart of a single-domain generalization method based on texture enhancement and morphological enhancement according to Embodiment 1 of the present invention; Figure 2 is a structural diagram of the texture enhancement module according to Embodiment 1 of the present invention; Figure 3 is a structural diagram of the morphological enhancement module according to Embodiment 1 of the present invention; Figure 4 is a training flowchart of the texture enhancement module according to Embodiment 1 of the present invention; Figure 5 is a training flowchart of the morphological enhancement module according to Embodiment 1 of the present invention; Figure 6 is an experimental result diagram according to Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The following combines the drawings and embodiments to further describe the specific embodiments of the present invention in detail. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0022] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0023] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0024] In addition, in the description of the present invention, unless otherwise stated, the meaning of "a plurality of" is two or more.

[0025] Embodiment 1 As Figure 1 shown, the flowchart of a single-domain generalization method based on texture enhancement and morphological enhancement according to a preferred embodiment of the present invention includes the following steps: S1: Obtain a first source domain image and a second source domain image, and perform preprocessing; In a feasible embodiment, the preprocessing includes image processing such as normalization, rotation, scaling, translation, etc., to ensure that the images have a consistent format and size before being input into each module.

[0026] S2: Input the preprocessed first source domain image into a trained texture enhancement module to obtain a corresponding first virtual sample; In a feasible embodiment, the structure of the texture enhancement module is shown in Figure 2 , and the loss functions used for training the texture enhancement module are the maximum entropy loss function and the minimum cross-entropy loss function, where the maximum entropy loss function is determined by the following formula:

[0027] where , , are the probabilities of predicting the occurrence of the i-th category event in the first source domain image, the first virtual sample, and the second virtual sample respectively, and i is the index of the category event; by maximizing the entropy loss, the Task Model (classification model) is encouraged to learn more general features instead of overfitting to specific samples in the source domain data. Therefore, this method helps the enhancement module to generate more diverse virtual samples, making it more capable of enriching the sample diversity of the source domain data. The minimum cross-entropy loss function is determined by the following formula:

[0028] where , , They are the predicted outputs of the first source domain image, the first virtual sample, and the second virtual sample respectively. , , are the class labels of the corresponding images, , , are the predicted classes of the classification model , , respectively. The semantic information consistency between the virtual samples and the source domain samples is ensured by minimizing the cross-entropy loss.

[0029] According to Figure 2 , it can be known that the texture enhancement module includes: an encoder, an adaptive instance normalization layer, a fully connected layer, and a decoder. Specifically, the encoder is used to obtain texture features according to the preprocessed first source domain image . The texture features include at least one of color, illumination, shadow, edge direction, local contrast, and object surface attributes; the fully connected layer is used to map random Gaussian noise to generate the mean and variance of the style image. The calculation method is as follows:

[0030] where is the fully connected layer mapping function, n is the random Gaussian function, , are the mean and variance of the style image features; The adaptive instance normalization layer AdaIN (Adaptive Instance Normalization) is used to replace the mean and variance of the texture features with the mean and variance of the style image to obtain the second texture feature. The decoder is used to obtain the first virtual sample according to the second texture feature. The calculation method is as follows: = Dec( ∙ ) where Enc(∙) and Dec(∙) are the encoder and the decoder, (∙) and (∙) respectively represent the mean and variance of the texture features, is the first virtual sample obtained after texture enhancement.

[0031] S3: Input the preprocessed first source domain image and the preprocessed second source domain image into the trained morphological enhancement module to obtain the corresponding second virtual sample; In a feasible embodiment, the structure of the morphological enhancement module is shown in Figure 3As shown, the loss function for training the morphological enhancement module is the same as that for training the texture enhancement module, which will not be elaborated here. Moreover, the processes for training the texture enhancement module and the morphological enhancement module are shown in Figure 4 , when training the texture enhancement module and the morphological enhancement module, the parameters of the classification model are fixed. Then, by maximizing the entropy loss and minimizing the cross-entropy loss, the generators for texture enhancement and morphological enhancement in the data enhancement module are optimized, thereby obtaining virtual samples with diverse texture styles and diversified morphologies, enriching the diversity of the training sample data. According to Figure 3 it is known that the morphological enhancement module includes a multi-scale morphological coordinate structure and a deformator. The multi-scale morphological coordinate structure is used to obtain a set of morphological key point coordinates based on the preprocessed first source domain image and the preprocessed second source domain image. Specifically, (1). The preprocessed first source domain image and the preprocessed second source domain image are input into an encoder with n feature extraction layers for feature extraction. The feature maps extracted by the nth feature extraction layer of the encoder are denoted as 、 , for 、 , the corresponding pixels are compared for similarity. The pixel coordinates with similarity greater than the threshold are denoted as the set of similar coordinates between the two feature maps of the nth feature extraction layer and the set of morphological coordinates . Among them, the similarity comparison is to calculate the squared difference between the pixel coordinates at the corresponding positions of the two feature maps. The calculation formula is as follows:

[0032] where, 、 are the pixel coordinates of the two feature maps 、 respectively. The pixel coordinates with the squared difference greater than the threshold are denoted as the set of similar coordinates between the two feature maps of the nth feature extraction layer and the set of morphological coordinates .

[0033] (2). The set of morphological coordinates 、 of the two feature maps at the nth layer are upsampled and mapped to obtain a coordinate set 、 of the same size as the two feature maps at the (n - 1)th layer. In this embodiment, the method used for upsampling is bilinear interpolation. Optionally, other methods such as nearest neighbor interpolation can also be used; during the bilinear interpolation process, specifically: , where is coordinates, then coordinates, and is the four adjacent coordinates; Dot product weighting: where are respectively , , coordinates; (3). Compare the corresponding pixels of the two feature maps , of the (n - 1)-th layer, and record the pixel coordinates with similarity greater than the threshold as the similarity coordinate set between the two feature maps of the (n - 1)-th layer; In this embodiment, the similarity comparison is to calculate the square difference between the pixel coordinates at the corresponding positions of the two feature maps, and the calculation formula is as follows:

[0034] where, , are respectively the pixel coordinates of the two feature maps , . At the same time, the calculation formula of the threshold is as follows:

[0035] where, is the mean value of the feature map , is the variance of the feature map . The calculation formula of

[0036] The calculation formula of

[0037] (4). Perform dot product weighted aggregation on the similarity coordinate set and the coordinate set to obtain the morphological coordinate set between the two feature maps of the (n - 1)-th layer; (5). Repeat steps (2), (3), and (4) until the morphological coordinate set of the two feature maps of the first layer feature extraction layer is obtained, for the morphological coordinate set of the first layer Process to obtain a set of morphological key point coordinates of the same size as the first source domain image and the second source domain image , where j represents the category of the source domain image, and at the same time the source domain image is only used for comparison and reference, while P represents the set of morphological key point coordinates of the source domain image . Since the morphological coordinate query is multi-scale, the morphological key point coordinates of the category object in the source domain image can be accurately found according to the feature context relationship, which can make the morphology of the category object in the virtual sample generated during the subsequent deformation of the source domain image conform to the actual situation

[0038] Furthermore, the deformer is used to obtain a second virtual sample according to the set of morphological key point coordinates and the first source domain image. Specifically, according to the set of morphological key point coordinates perform random deformation on the first source domain image to obtain an initially deformed image. Specifically, assume that the set of morphological key point coordinates has a total of key point coordinates, that is , randomly generate a displacement vector set with vectors , and translate the set of key point coordinates based on the displacement vector set to obtain an initially deformed image Generate a dense flow field according to the pixel coordinates of the first source domain image and the pixel coordinates of the deformed image, and determine parameters based on the dense flow field. The parameters include weight coefficients, polynomials constructed to solve linear equations, and augmented vectors Establish an inverse mapping function based on the parameters. The inverse mapping function is determined by the following formula

[0039] where , , are the weight coefficient, the polynomial constructed to solve the linear equation, and the augmented vector respectively represents the position of the pixel in the deformed image is the radial basis function matrix, specifically .

[0040] Generate the color of each pixel through bilinear sampling of the inverse mapping function and the pixel coordinates of the first source domain image, and output the final deformed image, that is, the second virtual sample. Next, this step is described: through and the relevant 2D displacement vector to specify deformation For each coordinate of the morphological key point Specify the target coordinates , and then use thin plate spline interpolation to interpolate the coordinates of the first source domain image to the deformed image to generate a dense flow field. This is a closed-form process that finds the parameters to minimize the affected by curvature constraints. With these parameters, the inverse mapping function can be obtained:

[0041] where represents the position of the pixel in the deformed image, gives the inverse mapping of the pixel in the source domain image, that is, the pixel coordinates in the source domain image, from which we can derive the color of the pixel , and then the color of each pixel can be generated by bilinear sampling, so as to output the deformed image .

[0042] S4: Use the first source domain image, the first virtual sample and the second virtual sample to train a classification model, and design a loss function to optimize the training process. When the number of training times reaches the specified number, stop training to obtain a trained classification model; In a feasible embodiment, the loss function for training the classification model is the same as the loss function for training the texture enhancement module, which will not be elaborated here. The process of training the classification model can be seen in Figure 5 .

[0043] S5: Input the target domain image into the trained classification model and output the classification result.

[0044] In this embodiment, by using the trained texture enhancement module and the trained morphological enhancement module to obtain the first virtual sample and the second virtual sample respectively, the sample domain of the generated virtual samples is enriched. Then, by jointly training the classification model with the first virtual sample and the second virtual sample, the data presentation method is enriched from multiple angles, which can better adapt to the changes of various test data sets, expand the receptive field of the classification model training, and achieve the purpose of improving the generalization of the model.

[0045] Embodiment 2 This embodiment is an experiment conducted to verify the effectiveness and superiority of the method proposed in Embodiment 1. Specifically, the PACS dataset is selected, which contains 4 different domains, namely photos P (1670 images), art paintings A (2048 images), cartoons C (2344 images), and sketches S (3929 images). Each domain contains 7 categories, such as horses, giraffes, houses, guitars, etc. In this embodiment, the size of the images is set to 224×224, the batch size of the training data is 32, the number of training epochs is 30, the Adam optimizer is used, and the training learning rate is 0.05. To ensure the fairness of model training, in this embodiment, the parameters of all comparison methods (baseline, PDEN, RandConv, L2D, ABA, and MetaCausal) and the method proposed in the present invention are set uniformly. Table 1 shows the experimental results obtained with the domain listed in the table header as the source domain and the remaining three domains as the target domains: Table 1: Model classification accuracy of different single-domain generalization algorithms under PACS

[0046] From the results in Table 1, it can be seen that the average accuracy based on the combination of texture enhancement and morphological enhancement is higher than that based on texture enhancement alone. This fully demonstrates the supplementary role of the morphological enhancement module in the texture enhancement module of the method of the present invention, and further illustrates that the method based on texture enhancement and morphological enhancement in Embodiment 1 of the present invention can better improve the generalization ability of the classification model.

[0047] Such as Figure 6 shown, when this embodiment conducts the second experiment on the PACS dataset, with the sketch domain as the source domain for the training data and the target domain, the training data for training the classification model are respectively sketch domain images, virtual samples obtained by texture enhancement, and virtual samples obtained by morphological enhancement. They are jointly input into the classification model to train a model with strong generalization ability. The target domain is the remaining three domains of the PACS dataset, namely photo domain images, art painting images, and cartoon images. According to Figure 6 it can be known that it can be clearly observed that there are obvious differences in the texture styles and morphologies of their category objects. Therefore, the method of the present invention for generating virtual samples with diverse texture styles and diversified morphologies based on texture enhancement and morphological enhancement can solve the domain difference problems caused by texture style differences and morphological differences.

[0048] Embodiment 3 The embodiment of the present invention also provides a computer-readable storage medium, on which a program of a single-domain generalization method based on texture enhancement and morphological enhancement is stored. When the program is executed, the steps of the described single-domain generalization method based on texture enhancement and morphological enhancement are implemented.

[0049] In summary, the embodiment of the present invention provides a single-domain generalization method and a storage medium based on texture enhancement and morphological enhancement. By using the trained texture enhancement module and the trained morphological enhancement module to obtain the first virtual sample and the second virtual sample respectively, the sample domain of the generated virtual samples is enriched. Then, by combining the first virtual sample and the second virtual sample to train the classification model, the data presentation method is enriched from multiple perspectives, which can better adapt to the changes of various test data sets, expand the receptive field of the classification model training, and achieve the purpose of improving the model generalization ability.

[0050] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and replacements can be made, and these improvements and replacements should also be regarded as the protection scope of the present invention.

Claims

1. A single-domain generalization method based on texture enhancement and morphology enhancement, characterized in that: The method comprises the following steps: Acquire a first source domain image and a second source domain image, and perform preprocessing; Inputting the preprocessed first source domain image into the trained texture enhancement module to obtain a corresponding first virtual sample; Inputting the preprocessed first source domain image and the preprocessed second source domain image into the trained morphological enhancement module to obtain a corresponding second virtual sample; Using the first source domain image, the first virtual sample and the second virtual sample to train a classification model, and designing a loss function to optimize the training process, when the number of training times reaches a specified number, stopping the training, and obtaining a trained classification model; The target domain image is input into the trained classification model, and the classification result is output.

2. The single-domain generalization method based on texture enhancement and morphology enhancement according to claim 1, characterized in that: The preprocessing includes at least one of normalization, rotation, scaling and translation.

3. The single-domain generalization method based on texture enhancement and morphology enhancement according to claim 1, characterized in that: The preprocessed first source domain image is input into a trained texture enhancement module to obtain a corresponding first virtual sample, wherein the texture enhancement module includes: an encoder, an adaptive instantiation layer, a fully connected layer and a decoder. Specifically, the encoder is used to obtain texture features according to the input preprocessed first source domain image, and the texture features include at least one of color, illumination, shadow, edge direction, local contrast, and object surface properties; the fully connected layer is used to map random Gaussian noise to generate a mean and variance of a style image; the adaptive instantiation layer is used to replace the mean and variance of the texture features with the mean and variance of the style image to obtain a second texture feature, and the decoder is used to obtain the first virtual sample according to the second texture feature.

4. The single-domain generalization method based on texture enhancement and morphology enhancement according to claim 1, characterized in that: The preprocessed first source domain image and the preprocessed second source domain image are input into a morphological enhancement module to obtain a corresponding second virtual sample, wherein the morphological enhancement module includes a multi-scale morphological coordinate structure and a deformer, the multi-scale morphological coordinate structure is used to obtain a morphological key point coordinate set according to the preprocessed first source domain image and the preprocessed second source domain image; the deformer is used to obtain a second virtual sample according to the morphological key point coordinate set and the first source domain image.

5. The single-domain generalization method based on texture enhancement and morphology enhancement according to claim 4, characterized in that: The multi-scale morphological coordinate structure is used to obtain a morphological key point coordinate set according to the preprocessed first source domain image and the preprocessed second source domain image, specifically including: (1) Input the preprocessed first source domain image and the preprocessed second source domain image into an encoder having n layers of feature extraction layers for feature extraction, and record the feature map extracted by the nth feature extraction layer of the encoder as , ,right , The corresponding pixels are compared for similarity, and the pixel coordinates with similarity greater than the threshold are marked as the similar coordinate set between the two feature maps of the nth feature extraction layer and morphological coordinate sets ; (2) The two feature maps of the nth layer , The morphological coordinate set Through upsampling, the mapping obtains two feature maps with the n-1th layer , Coordinate sets of the same size ; (3) The two feature maps of the n-1th layer , The corresponding pixels are compared for similarity, and the pixel coordinates with similarity greater than the threshold are marked as the similar coordinate set between the two feature maps of the n-1th layer ; (4) The similar coordinate set and the coordinate set Perform point multiplication weighted aggregation to obtain the morphological coordinate set between the two feature maps of the n-1th layer ; (5) Repeat steps (2), (3), and (4) until the morphological coordinate set of the two feature maps of the first feature extraction layer is obtained. , the morphological coordinate set for layer 1 Processing is performed to obtain a morphological key point coordinate set of the same size as the first source domain image and the second source domain image .

6. The single-domain generalization method based on texture enhancement and morphology enhancement according to claim 5, characterized in that: The similarity comparison is to calculate the square difference between the pixel coordinates of the corresponding positions of the two feature maps, and the calculation formula is as follows: in, , There are two feature maps , The pixel coordinates of .

7. The single-domain generalization method based on texture enhancement and morphology enhancement according to claim 5, characterized in that: The threshold is determined by the following formula: in, The feature map The mean of The feature map The variance of .

8. The single-domain generalization method based on texture enhancement and morphology enhancement according to claim 5, characterized in that: The deformer is used to obtain a second virtual sample according to the morphological key point coordinate set and the first source domain image, and the specific process is as follows: According to the morphological key point coordinate set The first source domain image is randomly deformed to obtain an image after initial deformation. Specifically, assuming that the morphological key point coordinate set Total The key point coordinates are , randomly generated with The displacement vector set of vectors , translating the key point coordinate set based on the displacement vector set to obtain an image after initial deformation; generating a dense flow field according to the pixel coordinates of the first source domain image and the pixel coordinates of the deformed image, and determining parameters based on the dense flow field, the parameters including weight coefficients, polynomials constructed by solving a linear equation group, and augmented vectors; Establishing an inverse mapping function based on the parameters; The inverse mapping function and the pixel coordinates of the first source domain image are bilinearly sampled to generate the color of each pixel, and the final deformed image, that is, the second virtual sample, is output.

9. The single-domain generalization method based on texture enhancement and morphology enhancement according to claim 8, characterized in that: The inverse mapping function is determined by the following formula: in, , , are the weight coefficient, the polynomial constructed by solving the linear equations, and the augmented vector, respectively. represents the position of the pixel in the deformed image, is the radial basis function matrix, specifically .

10. The single-domain generalization method based on texture enhancement and morphology enhancement according to claim 1, characterized in that: The loss functions of training the texture enhancement module, training the morphology enhancement module and training the classification model are the same, including maximizing the entropy loss function and minimizing the cross entropy loss function, wherein the maximizing entropy loss function is determined by the following formula: in, , , are respectively the probability of the occurrence of the i-th category event predicted in the first source domain image, the first virtual sample, and the second virtual sample, where i is the index of the category event; the minimization cross entropy loss function is determined by the following formula: in, , , are the predicted outputs of the first source domain image, the first virtual sample, and the second virtual sample, respectively. , , is the category label of the corresponding image, , , The classification model predicts the category , , probability.