Method and system for enhancing accuracy of identification model across source domain and target domain

By introducing spatial enhancement and spectral hybridization (SASMU) systems, combined with spatial data enhancement (SA) and spectral hybridization (SMU) technologies, augmented synthetic data sets are generated, which solves the domain gap problem in cross-origin and target domain recognition in the prior art, improves the accuracy of the face recognition model and enhances privacy protection.

CN120220203APending Publication Date: 2025-06-27OTOBRITE ELECTRONICS INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411752733.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-25
Filing Date
2024-12-02
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing facial recognition technology faces challenges such as privacy issues, long-tail distribution, inconsistent image quality, noise labeling and lack of attribute annotation when achieving high-precision recognition, especially in the recognition of cross-original domains and target domains, which affects model accuracy.

Method used

The spatial enhancement and spectral mixing (SASMU) system was introduced, combining spatial data enhancement (SA) and spectral mixing (SMU) technology, and the computation was performed in the spatial and frequency domains, and the amplitude and phase components were extracted through Fourier transform, which separated high-frequency and low-frequency components, and used Gaussian filters and soft allocation mapping to establish enhanced amplitudes, generate enhanced synthetic data sets, and trained recognition models without using real face images.

Benefits of technology

Effectively narrow the gap between the synthetic domain and the real domain, improve the accuracy of cross-original domain and target domain recognition models, enhance privacy protection during data processing, and do not require real face images during the training stage, achieving the most advanced face recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220203A_ABST
    Figure CN120220203A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for enhancing the accuracy of a recognition model across a source domain and a target domain. The method comprises the following steps of: respectively extracting amplitude and phase components from a source domain data set and a target domain data set; separating a high-frequency component of the target domain data set amplitude component from a low-frequency component of the source domain data set amplitude component; establishing an enhanced amplitude in the frequency domain by incorporating high frequency components separated from the target domain data set into low frequency components separated from the source domain data set; generating an enhanced synthetic dataset from the enhanced amplitude and the phase component of the source domain dataset; and training a recognition model with the enhanced synthetic data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for enhancing the accuracy of recognition models. More specifically, the present invention is about a method and system for enhancing the accuracy of recognition models across source and target domains. Background Art

[0002] Face recognition technology has become increasingly important in various fields, including security and identification systems. Nevertheless, existing systems have encountered significant challenges in achieving high accuracy, especially when faced with changes in spatial and spectral characteristics. In recent years, due to decades of progress in face recognition technology, it has become easier to deploy powerful face recognition products. Cutting-edge technologies can effectively handle profile image verification and perform well in processing images in the wild. However, with the rapid rise of privacy concerns, mainstream research heavily relies on large-scale web-crawled datasets, which raises issues of privacy infringement. The community has attempted to address this dilemma by training face recognition models using synthetic data, but this effort has faced significant domain gap challenges and requires access to real images and identity labels for model fine-tuning.

[0003] With the development of deep learning technology, modern face recognition methods have made significant progress in performance, achieving verification accuracy of over 99.5% on the Labeled Faces in the Wild (LFW) dataset and recognition accuracy (TAR) of 97.70% at FAR = 1e-4 on the IJB-C dataset. In addition to these successes, researchers have also extended the capabilities of modern face recognition technology to special applications, such as recognizing masked faces under near-infrared illumination conditions. However, many of these methods rely on web-crawled datasets, such as MS1M, CASIA-Webface, and WebFace260M, and still face various challenges:

[0004] Privacy issues: Obtaining consent from all individuals in a large dataset is an extremely complex task, such as WebFace260M with 4 million people and over 260 million face images.

[0005] Long-tail distribution: Datasets exhibit significant variations in the number of images, poses, and expressions per person.

[0006] Image quality: Maintaining consistent image quality in large datasets is challenging.

[0007] Noisy labels: Web-crawled image datasets may have noisy label problems because social networks often automatically label face images, resulting in occasional mislabeling.

[0008] Lack of Attribute Annotations: Comprehensive annotations of facial attributes such as pose, age, expression, and lighting are usually not available.

[0009] The main challenge is the privacy issue surrounding the use of identifiable information. Although attempts have been made to mitigate privacy concerns by adding unidentifiable noise or random area masking to facial images, there is still a risk of real and identifiable images being made public. To thoroughly address the privacy issue, using synthetic data to train facial recognition models has become a viable solution. Thanks to the progress of generative models and computer graphics technology, it is now possible to use computing resources to generate realistic images. However, the domain gap remains a major obstacle, and previous efforts often resort to using real images and labels to bridge this gap, thus weakening the effect of privacy protection.

[0010] Therefore, there is an urgent need for a solution that can address the above challenges and at the same time improve the accuracy of facial recognition networks. Summary of the Invention

[0011] This section extracts and compiles certain features of the present invention. Other features will be provided in subsequent paragraphs. Its purpose is to cover various modifications and similar arrangements within the spirit and scope of the appended claims.

[0012] To improve the accuracy of face recognition networks and address the challenges mentioned above, the present invention introduces a Spatial Augmentation and Spectrum Mixing (SASMU) system. This system combines spatial data augmentation (SA) based on synthetic datasets with spectrum mixing (SMU) techniques, operating in both the spatial domain and the frequency domain. Additionally, it includes methods for dataset preparation and spectrum mixing to control, render, and align synthetic faces. The proposed spectrum mixing method aims to narrow the gap between the synthetic domain and the real domain and introduce subsequent dataset statistical analysis. Specifically, the present invention provides analysis results of various options (including grayscale and perspective operations) to analyze the impact and potential of spatial data augmentation (SA), and applies spectrum mixing (SMU) to reduce the gap between the synthetic domain and the real domain, without requiring real face images during training and achieving state-of-the-art face recognition performance without using any personally identifiable information.

[0013] On the one hand, the present invention provides a method for enhancing the accuracy of an identification model across a source domain and a target domain, including the following steps: extracting amplitude and phase components from a source domain dataset and a target domain dataset respectively; separating high-frequency components of the amplitude component of the target domain dataset from low-frequency components of the amplitude component of the source domain dataset; establishing an enhanced amplitude in the frequency domain by incorporating the separated high-frequency components from the target domain dataset into the separated low-frequency components from the source domain dataset; generating an enhanced synthetic dataset based on the enhanced amplitude and phase components of the source domain dataset; and training an identification model with the enhanced synthetic dataset.

[0014] According to an embodiment of the present invention, a Fourier transform is applied to the source domain dataset and the target domain dataset to extract the amplitude and phase components.

[0015] According to an embodiment of the present invention, frequency components of the source domain dataset and the target domain dataset are obtained via the application of a two-dimensional discrete Fourier transform.

[0016] According to an embodiment of the present invention, the high-frequency components are separated by a high-pass Gaussian filter, and the low-frequency components are separated by a low-pass Gaussian filter.

[0017] According to an embodiment of the present invention, the enhanced amplitude is established by combining the low-frequency components of the source domain dataset and the high-frequency components of the target domain dataset, where the low-frequency components are modified using a Gaussian mask, and the high-frequency components are subjected to a complementary operation (1 - Gaussian), subtracting each value of the Gaussian mask from 1.

[0018] According to an embodiment of the present invention, a Gaussian-based soft assignment mapping is used to combine the high-frequency components of the target domain dataset and the low-frequency components of the source domain dataset to narrow the domain gap between the source domain dataset and the target domain dataset.

[0019] According to an embodiment of the present invention, the phase components (containing data requiring privacy protection) of the target domain dataset are filtered out during the generation of the enhanced synthetic dataset.

[0020] According to an embodiment of the present invention, an inverse discrete Fourier transform (DFT) or an inverse fast Fourier transform (FFT) is applied to the enhanced amplitude and phase components of the source domain dataset to generate the enhanced synthetic dataset.

[0021] According to an embodiment of the present invention, the enhanced synthetic dataset contains labels encoded in the phase components of the source domain dataset.

[0022] According to an embodiment of the present invention, the enhanced synthetic dataset is desensitized before being provided to the identification model.

[0023] On the other hand, the present invention provides a system for enhancing the accuracy of an identification model across a source domain and a target domain, comprising: a database storing a source domain dataset and a target domain dataset; a processing unit connected to the database, extracting amplitude and phase components from the source domain dataset and the target domain dataset respectively, and separating the high-frequency component of the amplitude component of the target domain dataset from the low-frequency component of the amplitude component of the source domain dataset; an integration unit connected to the processing unit, establishing an enhanced amplitude in the frequency domain by incorporating the separated high-frequency component from the target domain dataset into the separated low-frequency component from the source domain dataset; a dataset generation unit connected to the integration unit, generating an enhanced synthetic dataset based on the enhanced amplitude and phase components of the source domain dataset; and an identification model connected to the dataset generation unit, training with the enhanced synthetic dataset provided by the dataset generation unit.

[0024] According to an embodiment of the present invention, Fourier transform is applied to the source domain dataset and the target domain dataset to extract the amplitude and phase components.

[0025] According to an embodiment of the present invention, the frequency components of the source domain dataset and the target domain dataset are obtained via the application of two-dimensional discrete Fourier transform.

[0026] According to an embodiment of the present invention, the high-frequency component is separated by a high-pass Gaussian filter, and the low-frequency component is separated by a low-pass Gaussian filter.

[0027] According to an embodiment of the present invention, the enhanced amplitude is established by combining the low-frequency component of the source domain dataset and the high-frequency component of the target domain dataset, the low-frequency component is modified using Gaussian masking, and the high-frequency component undergoes a complementary operation (1 - Gaussian), subtracting each value of the Gaussian masking from 1.

[0028] According to an embodiment of the present invention, a Gaussian-based soft assignment mapping is used to merge the high-frequency component of the target domain dataset and the low-frequency component of the source domain dataset to narrow the domain gap between the source domain dataset and the target domain dataset.

[0029] According to an embodiment of the present invention, the phase component (containing data requiring privacy protection) of the target domain dataset is filtered out during the generation of the enhanced synthetic dataset.

[0030] According to an embodiment of the present invention, inverse discrete Fourier transform (DFT) or inverse fast Fourier transform (FFT) is applied to the enhanced amplitude and phase components of the source domain dataset to generate the enhanced synthetic dataset.

[0031] According to an embodiment of the present invention, the enhanced synthetic dataset contains tags encoded in the phase components of the source domain dataset.

[0032] According to an embodiment of the present invention, the enhanced synthetic dataset is desensitized before being provided to the recognition model. Description of the Drawings

[0033] Figure 1 is a block diagram showing the main components of a system for enhancing the accuracy of a recognition model across source and target domains according to an embodiment of the present invention.

[0034] Figure 2 is a flowchart showing a method for enhancing the accuracy of the recognition model across source and target domains according to an embodiment of the present invention.

[0035] Figure 3 Provides a conceptual overview showing the enhanced synthetic dataset generated according to an embodiment of the present invention.

[0036] Figures 4A to 4D Presents a conceptual overview showing various methods for generating the enhanced synthetic dataset as compared to Figure 3 the method of

[0037] Figure 5 Depicts an example comparison result of the performance / average accuracy of various methods using Figure 3 and Figures 4A to 4D

[0038] Figure 6 is a visualization result of various methods using Figure 3 and Figures 4A to 4D , as well as the PSNR values representing the image quality and similarity between the original synthetic image and the enhanced synthetic dataset.

[0039] Description of the Reference Numerals: 100 - System; 101 - Database; 102 - Processing Unit; 103 - Integration Unit; 104 - Dataset Generation Unit; 105 - Recognition Model; 111 - Source Domain Dataset; 112 - Target Domain Dataset. Detailed Description of the Embodiments

[0040] Certain features of the present invention are specifically set forth and pointed out herein; other features will be provided in the following description. This section will attempt to cover various modifications and arrangements made to encompass the spirit and scope of the claims.

[0041] ​The applicant declares that almost all elements are presented in the singular in the description or the drawings, but their meaning is at least one and not limited to one. In other words, in one embodiment, the number of an element may be one, while in another embodiment the number of that element may be greater than one. "One" should not be construed in the text and the drawings as a limitation on quantity.

[0042] Recognition models are widely used in various applications, but their performance may be hindered when dealing with source and target domains with a domain gap. The present invention addresses this problem by introducing novel methods and systems for enhancing the accuracy of recognition models across these domains. The present invention provides systems and methods for enhancing the accuracy of recognition models across source and target domains. Specifically, the present invention relates to the fields of machine learning and recognition models, particularly for manipulating frequency domain components to improve accuracy across source and target domains. The recognition model is trained on synthetic data within the source domain and applied to real images within the target domain. Although this embodiment focuses on face recognition, it is important to note that the applicability of the present invention extends beyond this use case and can be used in various fields, including advanced driver assistance systems (ADAS) and various camera applications.

[0043] Figure 1 In [the figure], the main components within system 100 are used to enhance the accuracy of the recognition model across source and target domains, as in the embodiments of the present invention. As shown in the figure, system 100 includes the following complete components: database 101, processing unit 102, integration unit 103, dataset generation unit 104, recognition model 105. Database 101 stores a source domain dataset 111 and a target domain dataset 112. Processing unit 102 is connected to database 101 to facilitate extracting amplitude and phase components from source domain dataset 111 and target domain dataset 112 respectively. Furthermore, processing unit 102 is responsible for isolating high-frequency components from the amplitude components of target domain dataset 112 and isolating low-frequency components from the amplitude components of source domain dataset 111. Integration unit 103 is interconnected with processing unit 102 to synthesize an enhanced amplitude in the frequency domain. This is achieved by combining the high-frequency components extracted from target domain dataset 112 with the low-frequency components extracted from source domain dataset 111. Dataset generation unit 104 is connected to integration unit 103 and is responsible for generating an enhanced synthetic dataset based on the enhanced amplitude and phase components from source domain dataset 111. Recognition model 105 is connected to dataset generation unit 104 with the enhanced synthetic dataset provided by dataset generation unit 104.

[0044] For a better understanding of the present invention, please refer to Figure 2, This is a flowchart showing a method for enhancing the accuracy of an identification model across source and target domains according to an embodiment of the present invention. The method of the present invention can be summarized as follows: Step S01 involves extracting amplitude and phase components from the source domain dataset 111 and the target domain dataset 112 respectively. In step S02, high-frequency components are separated from the amplitude components of the target domain dataset 112, while low-frequency components are separated from the amplitude components of the source domain dataset 111. Step S03 combines the high-frequency components separated from the target domain dataset 112 with the low-frequency components separated from the source domain dataset 111 to establish an enhanced amplitude in the frequency domain. Step S04 generates an enhanced synthetic dataset based on the enhanced amplitude established in step S03 and the phase components of the source domain dataset 111. Finally, in step S05, the enhanced synthetic dataset generated in step S04 is used to train the identification model 105.

[0045] Regarding frequency component extraction, the Fourier transform is applied to extract the amplitude and components within the source domain dataset 111 and the target domain dataset 112. Specifically, the frequency components of the two datasets are obtained through the application of two-dimensional discrete Fourier transform. As for high-frequency and low-frequency separation, a high-pass Gaussian filter is used to separate high-frequency components from the amplitude components of the target domain dataset 112, while a low-pass Gaussian filter is used to separate low-frequency components from the amplitude components of the source domain dataset 111.

[0046] Regarding enhanced amplitude establishment, the low-frequency components of the source domain dataset 111 and the high-frequency components of the target domain dataset 112 are combined to establish an enhanced amplitude. The low-frequency components are modified using Gaussian masking, and the high-frequency components undergo a complementary operation (1 - Gaussian), subtracting each value of the Gaussian mask from 1. Using Gaussian-based soft assignment mapping, this combination narrows the domain gap between the source domain dataset and the target domain dataset. Regarding privacy protection, the phase components of the target domain dataset 112 (containing data that requires privacy protection) are filtered out during the generation of the enhanced synthetic dataset.

[0047] Regarding enhanced synthetic dataset generation, the inverse discrete Fourier transform (DFT) or inverse fast Fourier transform (FFT) is applied to the enhanced amplitude and phase components of the source domain dataset 111 to generate the enhanced synthetic dataset. The enhanced synthetic dataset contains labels encoded in the phase components of the source domain dataset 111. Considering desensitization processing, the enhanced synthetic dataset undergoes desensitization processing before being provided to the identification model 105.

[0048] The main objective of the present invention is to use a synthetic dataset for training to establish a privacy-first face recognition model. The present invention introduces a breakthrough data enhancement technique called spectral mixing (SMU) to address the domain gap between real datasets and synthetic datasets, such as Figure 3 . Different from other frequency domain mixing strategies, such as Figures 4A to 4D, using weighted sum operation or hard assignment masking, the present invention integrates the amplitude components of synthetic and real data through Gaussian-based soft assignment mapping and enhances high-frequency information, such as Figure 3 .

[0049] In this specific embodiment, several assumptions can be made selectively: 1) semantic content, especially identity information, is mainly encoded in the phase component; 2) injecting the amplitude information of real data into synthetic information improves the alignment with the distribution of the real data set; 3) enhancing high-frequency information is proven to be more effective than low-frequency information. Therefore, the synthetic data essentially captures the real low-frequency information but lacks complex high-frequency details.

[0050] To better understand the present invention, the following is an exemplary formula for obtaining the frequency components of an image x ∈ R M×N . It should be understood that this is only an example and the present invention is not limited thereto.

[0051]

[0052] where (m, n) represents the coordinates of the image pixel in the spatial domain; x(m, n) is the pixel value; (u, v) represents the spatial frequency coordinates in the frequency domain; F(x)(u, v) is the complex frequency value of the image x; e and j are the Euler number and the imaginary unit respectively. Therefore, F -1 (·) is the two-dimensional inverse discrete Fourier transform that converts the spectrum to the spatial domain. The following is Euler's formula:

[0053] e jθ = cos(θ) + j sin(θ)

[0054] According to the above formula, the image is decomposed into orthogonal sine and cosine functions, which respectively constitute the imaginary part and the real part of the frequency component F(x). Then, the amplitude and phase spectra of F(x)(u, v) are defined as:

[0055]

[0056] where R(x) and I(x) represent the real part and the imaginary part of F(x) respectively.

[0057] Furthermore, a Gaussian kernel is used to establish a soft assignment mapping, denoted as G. The soft assignment mapping is defined as follows:

[0058]

[0059] where D0 is a positive constant representing the cut-off frequency, and D0 2 is the distance between the point (u, v) in the frequency domain and the center of the frequency rectangle, that is:

[0060] D(u, v) = ((u - M / 2) 2 + (v - N / 2)2 ) 1 / 2 ,

[0061] where M and N represent the height and width of the frequency rectangle and the image, respectively.

[0062] According to this embodiment, the following formula is further applied to two randomly sampled images x syn and x real to generate an enhanced synthetic dataset:

[0063]

[0064] where ○ represents element-wise multiplication operation. The low-frequency information of the synthetic data is retained, and the high-frequency details from the amplitude component of the real image are incorporated. The resulting amplitude component is then combined with the phase component of x syn to obtain the final enhanced synthetic image x' syn .

[0065] In summary, this embodiment uses soft assignment mapping to merge the low-frequency elements of the synthetic image with the high-frequency elements of the real image, generating a more realistic enhanced synthetic image. Importantly, this method specifically utilizes the amplitude spectrum of the real image to capture the high-frequency components without adding label or identity information during the training phase. This unique method allows this technique to be applied to different image datasets without manual annotation or labeling, making it a versatile tool suitable for various computer vision applications.

[0066] To better understand the efficacy of the method proposed by the present invention compared to alternative methods, please refer to Figures 3 to 6 . Figure 3 provides a conceptual overview showing how to generate an enhanced synthetic dataset. Figures 4A to 4D provides a conceptual overview showing various methods for generating an enhanced synthetic dataset compared to the method shown in Figure 3 . Figure 6 is the visualization results of various methods using Figure 3 and Figures 4A to 4D and the peak signal-to-noise ratio (PSNR) values indicating the image quality and similarity between the original synthetic image and the enhanced synthetic dataset.

[0067] Figure 4A In [reference], the amplitude spectrum of the source image is directly replaced by the amplitude spectrum of the target image, which results in inconsistent phase and amplitude of the synthetic image. Figure 4B uses a frequency mask / square mask to exchange the low-frequency components of the source amplitude spectrum, resulting in a ringing effect on the enhanced image because the square mask is used as an ideal filter. Figure 4C Combines the two amplitude spectra through a weighted sum operation without considering that different frequencies have different importance and information, which will generate artifacts in the enhanced image. Figure 4DThe high-frequency components of the source amplitude spectrum are retained, and the low-frequency components of the source are merged with the low-frequency components of the target image.

[0068] However, their configurations result in adjustments being made to only a limited number of frequency points on the synthetic image, only resulting in changes in the image intensity within the spatial domain. In other words, expanding the hyperparameters of these methods may cause ringing effects, as Figure 5 shown by the results. Additionally, the peak signal-to-noise ratio (PSNR) values of these enhanced images are calculated, as Figure 6 . The research results emphasize that, compared with the Figures 4A to 4D techniques, the Figure 3 method shown is excellent in producing high-quality images that are very similar to the original synthetic images.

[0069] The present invention introduces a novel method and system for enhancing the accuracy of recognition models across source and target domains through the manipulation of frequency components. By applying Fourier transform, Gaussian filter, and soft assignment mapping, the system can effectively address the domain gap and retain sensitive data while generating an enhanced synthetic dataset. These advancements contribute to the field of machine learning and recognition models, providing improved performance in a wide range of applications. In summary, the present invention has the following advantages: enhancing the accuracy of recognition models across source and target domains by narrowing the domain gap; enhancing privacy protection during the data processing process; and the flexible application of inverse Fourier transform in dataset generation.

[0070] The system also introduces a powerful face recognition system that utilizes the synthetic dataset to address privacy issues. The proposed method strategically combines spatial data augmentation (SA) and spectral mixture (SMU) to enhance data variation and narrow the gap between the synthetic domain and the real domain. First, a comprehensive analysis of common data augmentations under various real-world conditions and color spaces (such as RGB / grayscale space) is conducted to determine the optimal combination for face recognition using the synthetic dataset. Second, the factors causing the domain gap between the real dataset and the synthetic dataset are explored. Spectral mixture (SMU) is a novel frequency-domain mixing method aimed at bridging this gap and improving recognition performance. It should be noted that only synthetic data and real images (without labels) are used in the training phase, without incorporating the data in the target dataset.

[0071] Even though the present invention has been described in detail by the above embodiments and can be variously modified by those of ordinary skill in the art, all such modifications do not depart from the scope as claimed in the appended claims.

Claims

1. A method for enhancing the accuracy of a recognition model across source and target domains, characterized in that: The following steps are involved: Extracting amplitude and phase components from source domain dataset and target domain dataset respectively; Separating a high frequency component of the amplitude component of the target domain data set from a low frequency component of the amplitude component of the source domain data set; establishing an enhanced amplitude in the frequency domain by merging high frequency components separated from the target domain dataset into low frequency components separated from the source domain dataset; generating an enhanced synthetic dataset based on the enhanced amplitude and phase components of the source domain dataset; and The recognition model is trained with the enhanced synthetic dataset.

2. The method according to claim 1, characterized in that A Fourier transform is applied to the source domain data set and the target domain data set to extract the amplitude and the phase components.

3. The method according to claim 1, characterized in that The frequency components of the source domain dataset and the target domain dataset are obtained via application of a two-dimensional discrete Fourier transform.

4. The method according to claim 1, characterized in that The high frequency components are separated by a high pass Gaussian filter, and the low frequency components are separated by a low pass Gaussian filter.

5. The method according to claim 1, characterized in that The enhanced amplitude is established by combining the low frequency component of the source domain dataset modified using a Gaussian mask and the high frequency component of the target domain dataset subjected to a complementary operation by subtracting each value of the Gaussian mask from 1.

6. The method according to claim 1, characterized in that A Gaussian-based soft allocation mapping is used to merge the high-frequency component of the target-domain dataset and the low-frequency component of the source-domain dataset to reduce a domain gap between the source-domain dataset and the target-domain dataset.

7. The method according to claim 1, characterized in that The phase component of the target domain dataset is filtered out during generation of the enhanced synthetic dataset.

8. The method according to claim 1, characterized in that An inverse discrete Fourier transform or an inverse fast Fourier transform is applied to the enhanced amplitude and phase components of the source domain dataset to produce the enhanced synthetic dataset.

9. The method according to claim 1, characterized in that The enhanced synthetic dataset contains labels encoded in the phase component of the source domain dataset.

10. The method according to claim 1, characterized in that The enhanced synthetic dataset is desensitized before being provided to the recognition model.

11. A system for enhancing the accuracy of a recognition model across source and target domains, characterized in that: include: A database storing source domain datasets and target domain datasets; a processing unit, connected to the database, extracting amplitude and phase components from the source domain dataset and the target domain dataset respectively, and separating a high frequency component of the amplitude component of the target domain dataset from a low frequency component of the amplitude component of the source domain dataset; an integration unit, connected to the processing unit, for establishing an enhanced amplitude in the frequency domain by incorporating the high frequency components separated from the target domain data set into the low frequency components separated from the source domain data set; a data set generating unit, connected to the integrating unit, generating an enhanced synthetic data set according to the enhanced amplitude and phase components of the source domain data set; and The recognition model is connected to the dataset generation unit and is trained with the enhanced synthetic dataset provided by the dataset generation unit.

12. The system according to claim 11, characterized in that A Fourier transform is applied to the source domain data set and the target domain data set to extract the amplitude and phase components.

13. The system according to claim 11, characterized in that Frequency components of the source domain dataset and the target domain dataset are obtained via application of a two-dimensional discrete Fourier transform.

14. The system according to claim 11, characterized in that The high frequency components are separated by a high pass Gaussian filter, and the low frequency components are separated by a low pass Gaussian filter.

15. The system of claim 11, wherein: The enhanced amplitude is established by combining the low frequency component of the source domain dataset modified using a Gaussian mask and the high frequency component of the target domain dataset subjected to a complementary operation by subtracting each value of the Gaussian mask from 1.

16. The system of claim 11, wherein: A Gaussian-based soft allocation mapping is used to merge the high-frequency component of the target-domain dataset and the low-frequency component of the source-domain dataset to reduce a domain gap between the source-domain dataset and the target-domain dataset.

17. The system of claim 11, wherein: The phase component of the target domain dataset is filtered out during generation of the enhanced synthetic dataset.

18. The system of claim 11, wherein: An inverse discrete Fourier transform or an inverse fast Fourier transform is applied to the enhanced amplitude and phase components of the source domain dataset to produce the enhanced synthetic dataset.

19. The system of claim 11, wherein: The enhanced synthetic dataset contains labels encoded in the phase component of the source domain dataset.

20. The system of claim 11, wherein: The enhanced synthetic dataset is desensitized before being provided to the recognition model.

Citation Information

Cited By

  • FFT-based unsupervised domain adaptive hyperspectral image classification method and system

    CN121545056A