System and method for generating universal fingerprints using a controllable diffusion model with multimodal conditions

The DDPM-based GenPrint model provides explicit control over identity and appearance in fingerprint generation, addressing limitations of existing methods by generating diverse and realistic fingerprints for improved recognition system training and evaluation.

WO2025212123A1PCT designated stage Publication Date: 2025-10-09BOARD OF TRUSTEES OPERATING MICHIGAN STATE UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/038187
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-03
Filing Date
2024-07-16
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing fingerprint generation methods lack explicit control over both the identity and appearance of generated fingerprints, limiting their utility for training and evaluation of fingerprint recognition systems.

Method used

A denoising diffusion probabilistic model (DDPM) is used for fingerprint generation with explicit control over identity and appearance, leveraging multimodal conditions such as text and image embeddings to generate diverse and realistic fingerprints.

Benefits of technology

GenPrint generates highly realistic and diverse synthetic fingerprints with improved recognition performance, allowing for zero-shot generation of novel fingerprint styles and enhancing the training and evaluation of fingerprint recognition systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024038187_09102025_PF_FP_ABST
    Figure US2024038187_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A method for generating fingerprints includes receiving a plurality of fingerprint appearance factor signals, communicating the appearance factor signals to a diffusion model, and generating at the diffusion model a reference image based on the fingerprint appearance factor signals. The reference image includes a reference identifier. The method further includes providing, to a trained impression generator, training images with style embeddings, text embeddings or both. The method further includes generating a plurality of impressions comprising variations of the reference image formed by varying the style embeddings, the text embeddings, or both. The plurality of impressions includes an identifier corresponding to the reference identifier.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR GENERATING UNIVERSAL FINGERPRINTS USING A CONTROLLABLE DIFFUSION MODEL WITH MULTIMODAL CONDITIONSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 573,539, filed on April 3, 2024. The entire disclosure of the above application is incorporated herein by reference.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] This invention was made with government support under 17STCIN00001 awarded by the Department of Homeland Security via The Criminal Investigations and Network Analysis Center (CINA) at George Mason University. The government has certain rights in the invention.FIELD

[0003] The present disclosure relates to fingerprint recognition systems and, more particularly, to a method and system for generating synthetic fingerprints and training a fingerprint recognition system using the same.BACKGROUND

[0004] This section provides background information related to the present disclosure which is not necessarily prior art.

[0005] Fingerprint recognition systems have become ubiquitous and can be found in a plethora of different domains, such as forensics, healthcare, mobile device security, mobile payments, border crossing, and national identification. The systems rely on models that are trained with fingerprints typically from known databases. Recently, artificially generated fingerprints have been used during the training process.

[0006] The use of Artificial Intelligence Generated Content (AIGC) over the last few years has exploded due to advancements in model architectures and larger computers and data being used to train Generative Al (GAI) models. In particular, text generation models, such as ChatGPT, have catapulted the field of GAI into the publicview since its public release in November 2022. Following in its wake came stunning advancements in image and video generation models, such as ImageGen and SORA, utilizing denoising diffusion probabilistic model (DDPM) frameworks. Since their inception, DDPM models have proliferated themselves as the center of attention in many top computer vision conferences and journals within the last two years. Notably, their probabilistic framework and straight forward optimization process makes DDPMs more stable and easier to train compared to generative adversarial networks (GANs), one of the predominant frameworks for image generation previously. Furthermore, the advantages of diffusion models over GANs for image generation in terms of image quality have been demonstrated. Indeed, the recent surge in DDPM models has revolutionized GenAI capabilities across enumerable industries and applications.

[0007] Artificial fingerprint generation is one application which has shown increased interest in using synthetic data for training and evaluation of algorithms, aided by recent privacy and ethical concerns as well as difficulty and cost associated with collecting biometric data. Before the explosion of deep learning techniques, fingerprint generation methods began with intelligent, hand-crafted methods to simulate convincing fingerprint patterns and textures. Importantly, these methods allowed for generating multiple images of the same finger, opening the door to training and evaluation of fingerprint recognition algorithms.

[0008] Early GAN-based methods drastically improved the realism of the generated prints but lacked control over the fingerprint identity being generated. Subsequent works aimed to fill this gap by replacing each stage of the multi-stage generation pipeline of hand-crafted methods with GANs, preserving the identity of the generated fingerprints at each stage. However, besides identity, other appearance factors remained obscured and uncontrollable, such as the specific fingerprint class (e.g., whorl, left loop, right loop, plain arch, and tented arch), acquisition type (e.g., rolled, slap, contactless, swipe, and latent), sensor characteristics (e.g., optical, capacitive, thermal, etc.), and quality level (e.g., high, average, and low) of the generated prints. FPGAN- Control has been proposed to disentangle identity and appearance factors in the latent space and allowed for swapping between different appearance latent vectors to achieve some degree of control over intra-class variations (e.g., acquisition type, sensor, and pressure level); however, this method lacks explicit, humanly explainable control over appearance factors.

[0009] Recent advancements in text to image generation models utilizing DDPMs have demonstrated very realistic and controlled image generation capabilities.

[0010] One work utilized a combination of an elliptical shape generation model, mathematical ridge flow and Gabor filters for ridge pattern generation, and noise and distortion models to simulate realistic fingerprint patterns. Importantly, the model allowed for generating multiple impressions of the same finger leading to its adoption for aiding in training and evaluation of fingerprint recognition models. Despite its impressive capabilities and intelligent design, SFinGe is limited in its intra-class variations it is able to generate due to its hand-crafted nature. More recent methods have turned to deep learning techniques, starting with GANs, to learn the subtle intra-class variations that have led to more varied and realistic fingerprint images.

[0011] The introduction of GANs gave way to more realistic fingerprint generation that captured more realistic texture characteristics that are difficult to hand-design; however, early uses of GANs lacked control over the fingerprint identity being generated - severely limiting the utility of the generated fingerprints. One system used to try to fill this gap adopts CycleGAN as a wrapper around SFinGe generated images to impart them with more realistic textures, while leveraging SFinGe’s ability to generate multiple impressions. However, the intra-class and inter-class variations were still limited by the hand-designed generation of SFinGe. Another system went one step further and designed a multi-stage GAN method for generating highly realistic fingerprint ridge patterns with multiple impressions per finger and showed substantial improvement over SFinGe in utility for recognition model training. Finally, another system adopted a mixed variational autoencoder (VAE) and GAN architecture called FPGAN-Control to interpolate between latent identity and appearance vectors to be able to render fingerprint images in multiple different appearances. Still, the model lacked explicit control over the appearance factors and the possible space of generated fingerprint styles is constrained to the distribution of styles belonging to the original training set.

[0012] DDPMs have only just begun to be investigated for artificial fingerprint synthesis. A vanilla DDPM has been used to synthesize unconditional fingerprint patches and validated the realism compared to real fingerprint patches using the Frechet Inception Distance (FID) metric. Another applied an unconditional DDPM model trained on a dataset of latent, rolled, and plain (i.e., slap) fingerprint images to randomly generate fingerprint impressions of these types. The realism of the DDPM generated fingerprint images both in terms of NFIQ quantitative values and t-SNE qualitative plot comparisonsto the real fingerprint images. However, the model lacked control over both the identity and appearance of the generated fingerprints, which is critical to their utility for training and evaluation of fingerprint recognition models.SUMMARY

[0013] This section provides a general summary of the disclosure, and is not a comprehensive disclosure of its full scope or all of its features.

[0014] In the present disclosure, the proposed GenPrint model uses a denoising diffusion probabilistic model (DDPM) for fingerprint generation with explicit control over both the identity and appearance of generated fingerprint images. DDPM advancements are leveraged for controllable fingerprint image generation utilizing multimodal conditions (text and image) for improved generation capabilities. Text prompts are leveraged to allow for guidance of explainable appearance factors and rely on image style embeddings for factors not easily expressed in language. An added benefit of the novel image style condition set forth herein is that the generation outputs are no longer constrained to interpolating between the domain of the seen training data but allows for zero-shot generation of novel fingerprint sensor characteristics not seen during training.

[0015] Two stages are used for fingerprint generation. In the first stage, a finetuned stable diffusion model is used to generate full (i.e., rolled) fingerprint images of various fingerprint classes from a random noise vectors. For the first stage, the text prompt guiding the generation follows the template of “a rolled fingerprint image, {class} pattern”, high quality, ink on stock paper”, where the fingerprint class is randomly selected from the five available classes. This provides a full fingerprint ridge pattern for use in the subsequent generation stage which imparts controllable style variations to generate large intra-class variations. By varying the noise vectors for each generation, completely new and unique fingerprint patterns are generated. In the second stage, the generated fingerprint images from the first stage are passed through ID-Net and imparted with varying appearances based on the style embeddings (from reference images either belonging to the training set or from new example images from unseen sensors) and different text prompts providing explainable acquisition, sensor, and quality factors.

[0016] In one aspect of the disclosure, a method of generating fingerprints includes receiving a plurality of fingerprint appearance factor signals, communicating the appearance factor signals to a diffusion model, and generating at the diffusion model a reference image based on the fingerprint appearance factor signals. The referenceimage includes a reference identifier. The method further includes providing, to a trained impression generator, training images with style embeddings, text embeddings or both. The method further includes generating a plurality of impressions comprising variations of the reference image formed by varying the style embeddings, the text embeddings, or both. The plurality of impressions includes an identifier corresponding to the reference identifier.

[0017] Features of the method and system include receiving the plurality of fingerprint appearance factor signals by receiving a fingerprint class, an acquisition type, sensor characteristic, a quality level or combinations thereof; the fingerprint class comprises whorl, left loop, right loop, plain arch, or tented arch; the acquisition type comprising rolled, slap, contactless, swipe, or latent; the sensor characteristics comprising FTIR optical, direct-view optical, multispectral optical, capacitive or thermal; the quality level comprising high, average, and low; generating the reference image by generating the reference image from random noise vectors; prior to generating the plurality of impressions, masking the reference image; removing sensor dependent characteristics and style characteristics from the reference image to form a ridge pattern silhouette image to guide a spatial preservation of a fingerprint identity; prior to generating the plurality of impressions, masking the ridge pattern silhouette image; providing to the trained impression generator training images comprises providing to the trained impression generator training images with style embeddings; and providing to the trained impression generator training images comprises providing to the trained impression generator training images with text embeddings.

[0018] In another aspect of the disclosure, a fingerprint recognition system includes a noise generator generating random noise vectors, a user interface generating a plurality of fingerprint appearance factor signals, a reference identification image generator receiving a plurality of fingerprint appearance factor signals from the user interface, a reference image generator generating a reference image based on plurality of fingerprint appearance factor signals, said reference image comprising a reference identifier, and a trained impression generator trained with training images with style embeddings, text embeddings or both, said trained impression generator generating a plurality of impressions based on the reference image by varying style embeddings, the text embeddings, or both, said plurality of impressions comprising identifiers corresponding the reference identifier.

[0019] In another aspect of the disclosure, a fingerprint recognition system includes a noise generator generating random noise vectors, a user interface generating a plurality of fingerprint appearance factor signals, a reference identification image generator receiving a plurality of fingerprint appearance factor signals from the user interface, a reference image generator generating a reference image based on plurality of fingerprint appearance factor signals, said reference image comprising a reference identifier and a trained impression generator trained with training images with style embeddings, text embeddings or both, said trained impression generator generating a plurality of impressions based on the reference image by varying style embeddings, the text embeddings, or both, said plurality of impressions comprising identifiers corresponding the reference identifier.

[0020] The following features are set forth in the present disclosure.

[0021] GenPrint, a controllable latent diffusion model using text and image conditions for highly realistic and diverse synthetic fingerprint generation.

[0022] GenPrint is capable of generating fingerprints of any acquisition type, sensor, fingerprint class, and quality; including fingerprint styles not seen during training without any additional fine-tuning (e.g., zero-shot fingerprint style generation).

[0023] The generation process is controllable (both in appearance and identity preservation) and explainable via the use of humanly interpretable text prompts.

[0024] The realism of GenPrint generated images compared to prior fingerprint generator methods across various fingerprint realism metrics is improved.

[0025] The utility of GenPrint synthetic images is validated through experiments highlighting improved recognition performance of models trained on GenPrint images compared to real datasets and other benchmark fingerprint generation methods.

[0026] The utility of GenPrint images for evaluating fingerprint recognition systems by replacing real data for large-scale identification experiments is demonstrated.

[0027] Examples of synthetic images generated from SFinGe, PrintsGAN, FPGAN-Control and GenPrint were compared. Improved diversity of GenPrint was observed.

[0028] Further areas of applicability will become apparent from the description provided herein. The description and specific examples in this summary are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure.DRAWINGS

[0029] The drawings described herein are for illustrative purposes only of selected embodiments and not all possible implementations, and are not intended to limit the scope of the present disclosure.

[0030] Fig. 1 is a block diagrammatic view of the fingerprint generation system according to the present disclosure.

[0031] Fig. 2 is a block diagrammatic view of details of the GenPrint system of Fig. 1.

[0032] Fig. 3 is a flowchart of a method for operating the fingerprint generation system.

[0033] Fig. 4 is a chart showing training sets used for the fingerprint generation system.DETAILED DESCRIPTION

[0034] Example embodiments will now be described more fully with reference to the accompanying drawings.

[0035] The present disclosure sets forth a fingerprint generation system referred to as GenPrint. GenPrint is a multimodal latent diffusion model finetuned for fingerprint generation from a pretrained Stable Diffusion v1 .5 model with weights made available from the Diffusers library. The text to image fingerprint generation capabilities is described including the dataset curation process and fine-tuning procedure. The architectural design for incorporating style image embeddings into the Stable Diffusion pipeline is set forth. Also, the zero-shot style generation capability it facilitates is also set forth. Finally, the identity preservation process is described along with a detailed description of the full pipeline for generating synthetic fingerprint images with GenPrint.

[0036] Referring now to Figure 1 , a fingerprint generation system 10 is used for generating a plurality of fingerprint impressions or digital images and or digital representation of fingerprint impressions. Although the word impression is used, the generated fingerprints may not be specifically “impressed” onto a surface. In this disclosure, a digital image of a finger may also be called an impression. The controller 12 has a microprocessor 18 and a memory 20. The microprocessor 18 may be referred to as a processor. The memory 20 may be a non-transitory memory. The non-transitory memory 20 is a computer-readable medium that includes machine readable instructions that are executable by the processor 18. The machine readable instructions are used forgenerating fingerprints as described in great detail below. The controller 12 has the GenPrint system 30 disposed therein. Details of the GenPrint system 30 are set forth below.

[0037] The controller 12 is coupled to a style image bank 32, a diffusers library 34 and a noise generator 36. The noise generator 36 generates noise vectors that are used in the diffuser model.

[0038] Referring now to Figures 2, 3 and 4, the first step in fine-tuning stable diffusion for text to fingerprint generation is obtaining a large corpus of fingerprint images and associated text descriptions. A style image bank 32 provides training images and step 310 obtains the fingerprint images. For this purpose, aggregated data from multiple fingerprint datasets from predominately publicly available sources may be used as the style image bank 32. The training datasets are listed in Fig. 4 along with the acquisition label, sensor label, and number of images for each dataset. The aggregated dataset consists of data from five different acquisition types (rolled, slap, swipe, contactless, and latent) and thirty different sensing devices ranging from optical readers, capacitive, thermal, contactless, and even latent surfaces.

[0039] Missing from many fingerprint datasets are annotations for fingerprint class (whorl, plain arch, tented arch, left loop, and right loop) and quality (low, average, and high quality), which are needed to impart the generator with this kind of control. To obtain these labels, a style encoder 210 is utilized to obtain style embeddings 212 at step 312. The style embeddings 212 are ultimately communicated to a decoupled cross-attention block 214. Verifinger SDK v12.4 was used to extract class and NFIQ 2.0 quality estimations was used as the style encoder 210 in step 312. Since the NFIQ 2.0 metric was optimized for slap impressions utilizing frustrated total internal reflection (FTIR) optical imaging, the quality levels across each acquisition type may vary distinctly. Thus, quality distributions were fit according to a normal distribution using images belonging to each acquisition category and assigned low, average, and high quality labels to image clusters based on the mean ± standard deviation.

[0040] Using these annotations text prompt labels 220 are obtained in step 314 for each training image utilizing the following template: “a {acquisition} fingerprint image, {class} pattern, {quality} quality, {sensor}, {sensing}”, where the acquisition type is one of {rolled, slap, swipe, contactless, latent}, class is one of {whorl, plain arch, tented arch, left loop, right loop}, quality is one of {low, average, high}, sensor is one of the thirtytraining sensors listed Fig. 4, and sensing type is one of {FTIR optical, direct-view optical, multispectral optical, capacitive, thermal}.

[0041] For fine-tuning stable diffusion on the newly aggregated text to fingerprint image dataset, a low-rank adaptation (LoRA) strategy was used for more efficient training with a rank of 128. The LoRA weights are finetuned with a learning rate of 0.0001 , cosine scheduler, default Adam optimizer, and batch size of 96 spread across 8 Nvidia A100 GPUs. The model is trained for 500,000 steps and trained on fingerprint images of a resolution of 512x512 pixels.

[0042] Motivated by the fact that many of the textural intraclass variations present in fingerprint images are not easily expressed in language via simple text prompts, a deep learning-based representation was used to capture the characteristics. In particular, the pretrained VGG style encoder 212 was trained on ImageNet to embed style embeddings for each training image from the style image bank 32. The style embeddings 212 and text embeddings 232 are injected in step 316 into the diffusion model 240 via cross-attention layers 214 which are de-coupled from the cross-attention layers from the textual embeddings 232 used to control the explainable style factors received from a text encoder 234 based on the text prompt labels 220. That is, in step 318, style and text embeddings are coupled to the impression generator 230. The choice of VGG style embeddings 212 for style representation is motivated from two key insights: i.) the previous use of VGG for neural style transfer and ii.) visualizing the separation of VGG style embeddings for various fingerprint sensor types in the t-SNE embedding space.

[0043] During inference, style embeddings 212 from various sensor types present in the training data of the image bank 32 can be sampled to generate images for that sensor type. On the other-hand, even style embeddings extracted from images of a completely new, unseen sensor can be used to generate images in that new sensor domain. Therefore, the method is generalizable and allows for “zero-shot” fingerprint style generations without any additional fine-tuning required. This may be used to produce new fingerprint characteristics of latent, optical, capacitive, and contactless sensors outside those seen during training.

[0044] Several strategies for identity preservation and personalization in diffusion models have been proposed. Some of these techniques require additional fine-tuning for each new concept, whereas others can generate identity consistent generations for multiple subjects without inference time fine-tuning. To systems embed the identity of aninput reference image or images into the diffusion process via cross-attention layers. The guide for the diffusion model is to generate images which are identity consistent with the input reference images. Empirically, one system lacked the fine-grained spatial control needed to maintain the fingerprint ridge structure throughout the image. To solve this, ControlNet, which is another adaptation to the diffusion model process in which reference images are provided to the diffusion model, is used to guide the generation with spatially consistent outputs.

[0045] Leveraging the domain knowledge that the identity discriminative features of fingerprints which are consistent across multiple different acquisition and sensor types are the silhouettes of the ridge flow patterns giving rise to the relative orientation of minutiae points of each finger. ControlNet is therefore a suitable choice for imparting the DDPM model with identity preservation. Therefore, the ControlNet framework was used to provide explicit spatial consistency of the generated fingerprint ridge pattern by prepending a ridge extractor to the input of the identity preserving diffusion model, ID- Net. This ridge extractor 246 removes sensor dependent characteristics and other style characteristics from the Reference ID image 244 (input fingerprint control image) leaving only the ridge pattern silhouette image to guide a spatial preservation of the fingerprint identity, including the location and orientation of minutiae points. This combined with the text and style embeddings providing the style information, allows the ID-Net or trained impression generator 2308 to generate varying textural characteristics while maintaining the input fingerprint ridge pattern to form the multiple impressions 250. The architecture for the ridge extractor 246 is the light-weight SqueezellNet model, which has been successfully applied previously for fingerprint ridge extraction. The ridge pattern silhouette image is used by the impression generator 230 of the diffusion model.

[0046] The full generation pipeline for Gen Print consists of two stages. First, the fine-tuned stable diffusion block 240 has a reference identifier (ID) image generator 242 or reference image generator that is used to generate full (i.e., rolled) fingerprint images of various fingerprint classes from a random noise vector generated from the noise block 36. For this stage, fingerprint appearance factor signals are received from the user interface 14 at the reference ID image generator 242 in step 320. The signals may be received using guiding from prompts on the display 16 to form a generation that follows the template of “a rolled fingerprint image, {class} pattern”, high quality, ink on stock paper”, where the fingerprint class is randomly selected from the five available classes. A full fingerprint ridge pattern or reference ID image 244 has a reference identifier for usein the subsequent generation stage which imparts controllable style variations to generate large intra-class variations. By varying the noise vectors (random noise vectors) from the noise block 260 for each generation, completely new and unique fingerprint patterns or reference ID images 244 are generated in step 322.

[0047] In the second stage, the generated fingerprint images, the reference ID images 244 from the first stage are passed through ID-Net or trained reference generator 242 in step 324 that generates impressions 250 with varying appearances based on the style embeddings (from reference images either belonging to the training set or from new example images from unseen sensors) and different text prompts providing explainable acquisition, sensor, and quality factors.

[0048] Using the ControlNet framework was very successful at constraining the local spatial details to be preserved in the generated images; however, often the generated images had the tendency to over-constrain the generation process to preserve every detail of the input control image. This is undesirable if, for example, the desired output image is a slap fingerprint image where the input control image is a full, rolled fingerprint ridge pattern. The result is an unrealistic image with the full rolled fingerprint pattern in the style of the specified “slap” sensor input. A mask 252 is an optional feature to the output of the ridge extractor 246, when used, or the reference identifier (ID) image generator 242 and is used to aligned with the input text prompts and to apply a realistic foreground mask for the specified acquisition type in option step 326 prior to step 324. For example, if the prompt is to produce a slap fingerprint image, then an extracted mask of the fingerprint foreground area from one of the slap training images is applied to the input rolled fingerprint ridge pattern to produce an image with a fingerprint area resembling a realistic slap fingerprint. If instead the prompt is to generate a latent fingerprint, then a mask from a training latent fingerprint image is applied to the input control image or reference image to produce a realistic looking latent fingerprint with occluded areas of the ridge pattern.

[0049] Steps 310 through 324 may be repeated to generate a plurality of sets of fingerprints. Each set of fingerprints has a common reference identifier. As mentioned above, portions of fingerprints may be masked. The set of fingerprints may be used to trin or evaluate fingerprint recognition models in step 330.

[0050] Additionally, for the same reason as for including the masks, the ID-Net model or impression generator 230 may not produce realistic non-linear distortions to the generated images because it would modify the input fingerprint pattern supplied as theControlNet input. Therefore, realistic distortion grids may be randomly sampled to apply to the ControlNet image for each generation. These realistic distortion grids are obtained by computing minutiae displacements between genuine fingerprint pairs within the training dataset. During inference, an example distortion grid, indexed by the specified fingerprint acquisition type, is sampled and applied to the input reference image.

[0051] Different explicit control factors were evaluated for Gen Print to determine whether various text prompts are accommodated; namely, control over the fingerprint class, acquisition, sensor, and quality level.

[0052] Fingerprint Class: GenPrint is able to generate fingerprints of any of the five major classes of fingers: whorl, left loop, right loop, plain arch, and tented arch. Examples of each of the categories generated by GenPrint. The consistency of GenPrint generated images in following the fingerprint class prompt provided by the user is validated quantitatively using the commercially available fingerprint recognition software, Verifinger SDK v12.4. Specifically, 100 unique finger identities were generated using GenPrint in each of the five different fingerprint classes and classify each of the fingerprints using Verifinger and compute the accuracy between the Verifinger predictions and the ground truth class assigned by the input text prompts. The classification accuracy for whorl, left loop, and right loop fingerprints was 99%, indicating that 99 out of 100 generated fingerprints were classified by Verifinger as the same class intended to be generated by GenPrint. The classification accuracy for Verifinger on the plain arch (92%) and tented arch types (25%) was much more challenging for Verifinger, which often misclassified the arch type as either left or right loop in all the misclassifications. Understandably, these two fingerprint classes can be difficult to distinguish given the similarity in the ridge patterns.

[0053] Fingerprint Acquisition and Sensor Type: GenPrint was trained on data from 30 different acquisition devices which consists of various rolled, slap, swipe, contactless, and latent fingerprint acquisition types. Comparing GenPrint images and corresponding real images in the same sensor and acquisition types highlighted the realism and diversity in the possible generation space of GenPrint.

[0054] The results show a clear separation between very distinct acquisition devices and some small overlap in similar sensors such as the large number of different slap FTIR optical devices sharing similar characteristics. Furthermore, VGG style embeddings 212 were generated for 100 real fingerprint image examples in 5 different acquisition devices and embedded them into the t-SNE space along with theircorresponding generated images from Gen Print to show the similarity between corresponding real and synthetic images of the same acquisition device domains.

[0055] Quality Control: There are two ways in which GenPrint can manipulate the quality of the generated images. The first is through the text prompt where the user can specify either low, average, or high quality and the other is through passing a reference style image with a relatively low, average, or high quality appearance. Empirically, both approaches work well. For validating the quality control of GenPrint, 100 unique synthetic finger identities with 300 impressions of each of the five different acquisition types (rolled, slap, swipe, contactless, and latent) were generated and used the text prompt to generate 100 of those impressions for each quality level (low, average, and high). A clear separation was found among each of the quality levels across each of the acquisition types, verifying GenPrint’s appropriate control over the quality of the generated fingerprints.

[0056] To validate the quality of zero-shot fingerprint style generation, an experiment using t-SNE visualizations was performed using embed 100 example synthetic and real images from 6 different acquisition device domains from an unseen dataset which was not included in the training dataset for GenPrint. These images come from the recently released the LFIW dataset. Very close similarity to corresponding real and synthetic images of the same acquisition devices was observed, demonstrating GenPrint’s adaptability toward zero-shot style generation from novel acquisition devices.

[0057] One criteria for the quality of synthetic fingerprint generators is the utility for training better fingerprint recognition models. The utility of GenPrint both when training on only synthetically generated images and when augmenting a set of real fingerprint images with additional synthetic data was determined. For a baseline comparison, a comparison using several previous synthetic fingerprint generators was performed including SFinGe, PrintsGAN, and FPGAN-Control.

[0058] ResNet50 recognition models were trained using an ArcFace loss function on incremental subsets of each database using increments from 1 ,600 identities (the size of the real N2N fingerprint database) to 35,000 identities. Experimentally, improved results were found.

[0059] In addition to being useful for training, synthetic fingerprints can also help with large-scale evaluation of fingerprint recognition algorithms, where collecting a dataset of potentially millions of unique real fingers can be prohibitively expensive. To demonstrate the feasibility of GenPrint images to be used for such purposes a largedatabase of 64,000 unique rolled fingerprints were generated and compared with a database of 64,000 real rolled fingerprint identities from the MSP dataset as a background gallery for a latent to rolled fingerprint search using latent probes and corresponding mates from the NIST SD 27 latent dataset. Ideally, the search performance should be similar when using the real fingerprint background images and GenPrint fingerprint background images. The experiments were repeated using a database of 64,000 unique identities from FPGAN-Control as a baseline. The results showed better overlap in the search accuracies between GenPrint background gallery, and the real fingerprint gallery compared to the overlap between FPGAN-Control and the real dataset - indicating that GenPrint images make a more suitable replacement for real images for large-scale search evaluations than the baseline FPGAN-Control method. In particular, the rank-1 accuracy on the real background dataset is 82.17%, whereas it was 82.95% and 83.72%, for GenPrint and FPGAN-Control, respectively.

[0060] The present system provides a significant advancement in the field of synthetic fingerprint generation by addressing key limitations in current synthetic data generation methods. By employing latent diffusion models with multimodal conditions, GenPrint offers a versatile framework capable of generating diverse fingerprint images while preserving identity and providing humanly understandable control over various appearance factors. Unlike previous approaches, GenPrint is not constrained by the characteristics of the training dataset alone, allowing for the generation of novel sensor and style attributes during inference without the need for additional fine-tuning. The experimental results, based on a variety of publicly available datasets, highlight the efficacy of GenPrint in terms of identity preservation, explicit and explainable control, and narrowing the gap between synthetic and real domains. Moreover, the universality of Gen Print-generated images demonstrates comparable or even superior accuracy compared to models trained solely on real data, thus enhancing the performance and generalization of fingerprint recognition systems. Overall, the utilization of GenPrint presents promising implications for advancing fingerprint recognition technologies while addressing privacy concerns associated with sensitive biometric data.

[0061] Example embodiments are provided so that this disclosure will be thorough and will fully convey the scope to those who are skilled in the art. Numerous specific details are set forth such as examples of specific components, devices, and methods, to provide a thorough understanding of embodiments of the present disclosure. It will be apparent to those skilled in the art that specific details need not be employed, thatexample embodiments may be embodied in many different forms and that neither should be construed to limit the scope of the disclosure. In some example embodiments, well- known processes, well-known device structures, and well-known technologies are not described in detail.

[0062] The terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. As used herein, the singular forms "a,” "an," and "the" may be intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises," "comprising," “including,” and “having,” are inclusive and therefore specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order discussed or illustrated, unless specifically identified as an order of performance. It is also to be understood that additional or alternative steps may be employed.

[0063] When an element or layer is referred to as being "on," “engaged to,” "connected to," or "coupled to" another element or layer, it may be directly on, engaged, connected or coupled to the other element or layer, or intervening elements or layers may be present. In contrast, when an element is referred to as being "directly on," “directly engaged to,” "directly connected to," or "directly coupled to" another element or layer, there may be no intervening elements or layers present. Other words used to describe the relationship between elements should be interpreted in a like fashion (e.g., “between” versus “directly between,” “adjacent” versus “directly adjacent,” etc.). As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0064] Although the terms first, second, third, etc. may be used herein to describe various elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms may be only used to distinguish one element, component, region, layer or section from another region, layer or section. Terms such as “first,” “second,” and other numerical terms when used herein do not imply a sequence or order unless clearly indicated by the context. Thus, a first element, component, region, layer or section discussed below couldbe termed a second element, component, region, layer or section without departing from the teachings of the example embodiments.

[0065] Spatially relative terms, such as “inner,” “outer,” "beneath," "below," "lower," "above," "upper," and the like, may be used herein for ease of description to describe one element or feature's relationship to another element(s) or feature(s) as illustrated in the figures. Spatially relative terms may be intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is turned over, elements described as "below" or "beneath" other elements or features would then be oriented "above" the other elements or features. Thus, the example term "below" can encompass both an orientation of above and below. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly.

[0066] The foregoing description of the embodiments has been provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure. Individual elements or features of a particular embodiment are generally not limited to that particular embodiment, but, where applicable, are interchangeable and can be used in a selected embodiment, even if not specifically shown or described. The same may also be varied in many ways. Such variations are not to be regarded as a departure from the disclosure, and all such modifications are intended to be included within the scope of the disclosure.

Claims

What is claimed is:1 . A method for generating fingerprints comprising: receiving a plurality of fingerprint appearance factor signals; communicating the appearance factor signals to a diffusion model; generating at the diffusion model a reference image based on the fingerprint appearance factor signals, said reference image comprising a reference identifier; providing to a trained impression generator training images with style embeddings, text embeddings or both; and generating a plurality of impressions comprising variations of the reference image formed by varying the style embeddings, the text embeddings, or both, said plurality of impressions comprising an identifier corresponding the reference identifier.

2. The method of claim 1 wherein receiving the plurality of fingerprint appearance factor signals comprises receiving a fingerprint class, an acquisition type, sensor characteristic, a quality level or combinations thereof.

3. The method of claim 2 wherein the fingerprint class comprises whorl, left loop, right loop, plain arch, or tented arch,4. The method of claim 2 wherein the acquisition type comprises rolled, slap, contactless, swipe, or latent.

5. The method of claim 2 wherein the sensor characteristics comprise FTIR optical, direct-view optical, multispectral optical, capacitive or thermal.

6. The method of claim 2 wherein the quality level comprises high, average, and low.

7. The method of claim 1 wherein generating the reference image comprises generating the reference image from random noise vectors.

8. The method of claim 1 wherein prior to generating the plurality of impressions, masking the reference image.

9. The method of claim 1 further comprising removing sensor dependent characteristics and style characteristics from the reference image to form a ridge pattern silhouette image to guide a spatial preservation of a fingerprint identity.

10. The method of claim 9 wherein prior to generating the plurality of impressions, masking the ridge pattern silhouette image.

11. The method of claim 10 wherein providing to the trained impression generator training images comprises providing to the trained impression generator training images with style embeddings.

12. The method of claim 11 wherein providing to the trained impression generator training images comprises providing to the trained impression generator training images with text embeddings.

13. A fingerprint recognition system comprising: a noise generator generating random noise vectors; a user interface generating a plurality of fingerprint appearance factor signals; a reference identification image generator receiving a plurality of fingerprint appearance factor signals from the user interface; a reference image generator generating a reference image based on plurality of fingerprint appearance factor signals, said reference image comprising a reference identifier; and a trained impression generator trained with training images with style embeddings, text embeddings or both, said trained impression generator generating a plurality of impressions based on the reference image by varying style embeddings, the text embeddings, or both, said plurality of impressions comprising identifiers corresponding the reference identifier.

14. The system of claim 13 wherein the plurality of fingerprint appearance factor signals comprises a fingerprint class, an acquisition type, sensor characteristic, a quality level or combinations thereof.

15. The system of claim 14 wherein the fingerprint class comprises whorl, left loop, right loop, plain arch, or tented arch; the acquisition type comprises rolled, slap,contactless, swipe, or latent; the sensor characteristics comprise FTIR optical, direct- view optical, multispectral optical, capacitive or thermal and the quality level comprises high, average, or low.

16. The system of claim 13 wherein generating the reference image comprises generating the reference image from random noise vectors.

17. The system of claim 13 further comprising a mask masking the reference image.

18. The system of claim 13 further comprising a ridge extractor removing sensor dependent characteristics and style characteristics from the reference image to form a ridge pattern silhouette image to guide a spatial preservation of a fingerprint identity.

19. A fingerprint recognition system comprising: a noise generator generating random noise vectors; a user interface generating a plurality of fingerprint appearance factor signals; a reference identification image generator receiving a plurality of fingerprint appearance factor signals from the user interface; a reference image generator generating a reference image based on plurality of fingerprint appearance factor signals, said reference image comprising a reference identifier; and a trained impression generator trained with training images with style embeddings, text embeddings or both, said trained impression generator generating a plurality of impressions based on the reference image by varying style embeddings, the text embeddings, or both, said plurality of impressions comprising identifiers corresponding the reference identifier.

20. The fingerprint recognition system of claim 19 further comprising ridge extractor removing sensor dependencies and style characteristics from the reference image to form a ridge pattern silhouette image, said trained impression generator generating the plurality of impressions based on the ridge pattern silhouette image.

Citation Information

Patent Citations

  • Palm print image synthesis method and system

    CN116434006A