OCT fingerprint section image authenticity detection method based on reconstruction difference
Through the combination of a fully convolutional neural network model and feature extractor, the problem of noise and background information interference in the OCT fingerprint section image is solved, and efficient and accurate fake image detection is achieved, reducing sample training costs.
Patent Information
- Application Number
- CN202210191133.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-02-25
AI Technical Summary
The existing OCT-based fingerprint recognition system lacks efficient and accurate automatic identification methods, especially when processing forged OCT fingerprint section images, there is noise and background information interference, affecting the discrimination results.
The fully convolutional neural network model is used, and only real finger B-scan images are trained. The images are reconstructed through the encoder and generator, and the feature extractor is used to reduce background information interference, and authenticity is judged by reconstructing differences and feature similarity.
It realizes efficient and accurate detection of forged OCT fingerprint section images without a large number of complex preprocessing, reducing sample training costs and improving model generalization.
Smart Images

Figure CN114581963B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biometric identification and anomaly detection, and is specifically applied to detecting forged OCT fingerprint section images. Background Art
[0002] One of the major features of optical coherence tomography (OCT) is that it can detect two-dimensional or three-dimensional structural images of biological tissues. When applied to fingers, it can detect subcutaneous information of the fingers. In addition to being used to reconstruct and identify fingerprints, it also improves the ability to detect live bodies and has a certain degree of anti-counterfeiting ability. However, the current fingerprint recognition system based on OCT usually requires human participation to judge the authenticity of the image after the image is collected, and there is still a lack of an efficient and accurate automatic anti-counterfeiting method.
[0003] Fake sample detection is a specific application in anomaly detection. In recent years, related research in the field of anomaly detection has more frequently used deep learning methods, which are simpler to apply and have better detection performance than traditional methods. Although conventional neural network classification models have achieved good true-false discrimination, they require positive and negative samples when training the model, and the number of the two samples must be balanced. In addition, they are generally closed-loop models. If the amount of data increases, the accuracy will decrease, and they lack good generalization. These problems undoubtedly add a lot of training costs. Therefore, an idea of using only one type of data for training has emerged, which can also be called a single-category classification model. The purpose is to identify only the category involved in the training as much as possible, and the other categories are directly classified as negative categories. At present, generative models such as autoencoders and adversarial generative networks are commonly used to achieve category distinction based on the degree of reconstruction difference. However, the direct application of this network structure to the detection of forged OCT fingerprint cross-section images is not ideal. This is mainly because OCT fingerprint cross-section images are not natural images. If they are not well preprocessed, there will be a lot of irrelevant noise and background information, which will affect the final discrimination results. Summary of the invention
[0004] The present invention aims to overcome the shortcomings of the prior art in identifying forged OCT fingerprint section images and provide a simple, automated method for detecting forged fingerprint OCT section images that does not require a large amount of complex preprocessing.
[0005] The present invention is a part of the OCT fingerprint recognition system, belonging to the image quality judgment part, and aims to screen out forged fingerprints before fingerprint recognition, thereby improving recognition accuracy.
[0006] The basic implementation principle of the present invention is to train the neural network model using only positive samples (real finger B-scan images). Since the model is only trained in positive samples, the model self-learns the data distribution of positive samples, including the data distribution of both latent space and image pixels. The model only has a good reconstruction effect on positive samples. The image quality generated by this type of sample after the encoder and decoder is high, and the difference with the input image is not large, with a small reconstruction difference. However, if it is a negative sample (pseudo-body B-scan image), the above effect cannot be achieved. After the model training is completed, if the input image is a negative sample, since the restored image is more like the real image, it will show a large difference compared with the input image. According to this difference, a threshold can be set for authenticity discrimination.
[0007] Among them, since the noise existing in the input image will increase the difference in pixels between the two images after reconstruction, it is inaccurate to use only pixel differences as a criterion for measuring authenticity. The features extracted by the neural network can better reflect the main semantic information and solve this problem to a certain extent. Therefore, the difference comparison in the method of the present invention is mainly compared at the feature vector level.
[0008] The method proposed in the present invention is an OCT fingerprint section image counterfeit identification method based on reconstruction difference, and the specific steps are as follows:
[0009] Step S1, construct a fully convolutional neural network model. The main body of the model consists of three parts: an encoder, a generator, and a feature extractor, as shown in Figure 1. The encoder first obtains the feature map of the data distribution of the input image in the latent space, and then uses the generator to reconstruct an image similar to the input image from the obtained data distribution. Since similarity evaluation is required in the feature space, but the feature map information finally output by the decoder still has a high degree of feature coupling, and the background in the original image occupies a large part, which makes the feature retain a considerable part of the background information in the original image, and it is difficult to directly use it as a feature representation of the image for subsequent feature comparison. Therefore, a feature extractor, that is, a feature extraction module, is added to the model to obtain a more semantic feature representation of the input image. This part uses ResNet as the basic structure. In order to more accurately locate the area of interest in the image and reduce background content interference, a channel attention module and a spatial attention module are added.
[0010] Step S2: Prepare training data and test data. Collect images collected by the OCT system, of which B-scan images from real fingers of different individuals are used as positive sample images, and B-scan images from phantoms made of different imitation materials are used as negative sample images. In addition, 10 images of the OCT system without the object to be tested are collected, which are only background images. However, due to the poor quality of some original collected images, especially too much useless information on the left and right sides of the image, it is necessary to enhance the image before training. The specific process is: perform image cropping operation on the B-scan image with an original size of 1800*500, crop 200 pixels on the left and right of the original image, and obtain a 1400*500 B-scan image, and then adjust the image size, use the bicubic interpolation method, and scale the cropped image size to the required size. In the experiment, it is scaled to 256*256 and converted to a grayscale image. After the preprocessing method, only 70% of the positive sample images are randomly selected as training data from the positive sample images. Select another 30% of positive and negative sample images, and use them as test data after the number is balanced. Perform data enhancement on 10 images containing only background to expand the number to 100 and save them for subsequent operations. The specific methods of data enhancement include: random cropping and then resizing to the original size, random Gaussian blur, and random flipping.
[0011] Step S3: training the network model. The overall training process can be seen in the attached figure. Figure 2 . Select the divided training images as input data. Each time the data is loaded, the original image data is backed up and then randomly blocked with black blocks of random size. The image data obtained after the blocking operation is recorded as x′. It passes through the encoder E(*) and the generator G(*) to obtain the corresponding reconstructed image, recorded as G(E(x′)). Calculate the difference between the reconstructed image and the original unblocked input image in terms of pixels. The expected difference value is as small as possible so that the distribution of the generated image is as close as possible to the original input image. Use the L1 Loss mean absolute error, recorded as the reconstruction error L recon , calculated as follows:
[0012] L recon =||G(E(x′))-||x|| 1 (1)
[0013] Where x represents the data distribution of the original input image, and G(E(x′)) represents the data distribution of the image reconstructed and restored after the network model. This loss function is only applied to the encoder and generator parts to improve the image reconstruction quality.
[0014] In order to alleviate the overfitting problem that may exist in the later stage of feature extractor training and improve the robustness of the model, a simple data augmentation operation can be used to amplify the data. It is necessary to vertically flip x and G(E(x)) to obtain the corresponding enhanced image data x^ and G(E(x))^. A total of 4 sets of data, including unenhanced and enhanced data, are input into the feature extractor. The obtained feature vector is taken as the positive feature vector, denoted as z pos At the same time, the same number of enhanced background image data prepared in step S2 is randomly selected and sent to the feature extractor. The feature vector obtained in this part is used as a negative feature vector, denoted as z neg . Start with z pos Select a positive eigenvector as the anchor point, denoted by z o , and then paired with another feature vector in the same batch. Among these combinations, the pairing combination of the anchor point and the positive feature vector is called a positive data pair, and the pairing combination with the negative feature vector is called a negative data pair. Assuming that the total number of feature vectors is M, after the above combination operation, 3 groups of positive data pairs, M-4 groups of negative data pairs, and a total of M-1 groups can be obtained. Then select the remaining positive feature vectors in turn and repeat the above operation.
[0015] The goal is to expect positive data pairs to have high similarity, while negative data pairs to have low similarity. The similarity of two vectors in a data pair is reflected by the cosine similarity calculation. The closer the value is to 1, the more similar the two vectors are, as shown in the following formula:
[0016]
[0017] Among them, S(a, b) is represented by the vector z a With vector z b Cosine similarity of data pairs, * T represents vector transpose, ||*|| represents the modulus of the vector, and γ is the scale parameter used to adjust the original [-1, 1] range of cosine similarity.
[0018] After determining the similarity measurement standard, set the contrast loss function L con , this loss function is similar to the softmax-cross entropy loss function in definition. In the process of optimizing the loss function, the proportion of positive data pairs similarity is gradually increased, so as to achieve the learning goal of the feature extractor part: maximizing the similarity of positive data pairs and minimizing the similarity of negative data pairs. First, calculate the proportion of positive data pairs composed of one anchor point in all combinations containing the anchor point. The goal is to expect the proportion to be as large as possible, so the loss function needs to take a negative sign again, as shown in the following formula:
[0019]
[0020] Among them, L con_anchor_nrepresents the average loss value of the positive data pair with the nth positive eigenvector as the anchor point, and M is the average loss value of the positive data pair with the anchor point z o_n The total number of positive data pairs, S(z o_n , z pos_i ) represents the i-th anchor point z o_n The cosine similarity of the positive data pair, N is the anchor point z o_n The total number of negative data pairs, S(z o_n , z neg_j ) represents the jth containing anchor point z o_n The cosine similarity of the negative data pairs.
[0021] Next, the loss values of the remaining anchor point combinations are calculated, and the above calculations are performed in sequence. Finally, the loss values obtained by all anchor point combinations are summed and averaged to obtain the final contrast loss L of the feature extractor part. con .
[0022]
[0023] Where N is the total number of anchor points, and this loss function is only applied to the feature extractor part.
[0024] After setting the loss function, the constructed network model is trained for multiple rounds. The model weight parameters are updated and optimized through back propagation until the loss function tends to converge, at which time the training can be stopped.
[0025] Step S4: Test the network model. The overall test process can be seen in the attached figure. Figure 3 The divided test data is selected as the input data, denoted as x, for testing. The testing process is similar to the training process in step S3. x passes through the encoder E(*) and the generator G(*) to obtain the corresponding reconstructed image, denoted as G(E(x)), which is input into the feature extractor. The cosine similarity is also used to calculate the feature vector z corresponding to x and G(E(x)). 1 、z 2 The similarity is calculated and saved. Generally speaking, the similarity of positive samples is generally high, and that of negative samples is generally low. Then, the ROC curve is drawn based on the cosine similarity calculation results of all test data, and the appropriate threshold is set based on the accuracy, false detection rate, and missed detection rate. It is set that as long as the cosine similarity calculation is higher than the threshold, it can be identified as a real finger image, otherwise it is identified as a fake finger image.
[0026] The advantages of the present invention are:
[0027] Compared with the conventional neural network model used for anti-counterfeiting, which is a typical binary classification network model, the network model proposed in the present invention requires a balance between true and false data for model training. However, the present invention only needs to use one category of data and add a small amount of supplementary data. In practical applications, real finger B-scan images are used as the main training data, and B-scan background images are used as supplementary data, which effectively reduces the sample training cost of the network model.
[0028] In the feature extractor proposed by the network model of the present invention, the channel and spatial attention mechanisms are used to reduce the interference of background information on the proposed features to a certain extent. The contrast loss function is used to make positive samples closer and positive and negative samples farther apart, thereby improving the feature distinction of positive and negative samples and having good generalization.
[0029] The model proposed in the present invention is an end-to-end model that does not require complex preprocessing or the preservation of feature vectors of standard positive samples. By using the trained model and inputting a conventional OCT section image, the authenticity of the image can be determined based on the difference between the input image and the reconstructed image in the feature vector. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1a to Figure 1c is a structural diagram of the neural network model of the present invention, wherein Figure 1a is the encoder and generator network structure, Figure 1b is the feature extractor network structure diagram, Figure 1c Feature extractor attention module diagram;
[0031] Figure 2 It is a training flow chart of the detection model of the present invention;
[0032] Figure 3 It is a test flow chart of the detection model of the present invention;
[0033] Figure 4a to Figure 4b is the positive sample input in the experiment of the present invention and the image reconstructed by the corresponding model, where Figure 4a is the input image, Figure 4b is the reconstructed image;
[0034] Figure 5a to Figure 5b is the negative sample input in the experiment of the present invention and the image reconstructed by the corresponding model, where Figure 5a is the input image, Figure 5b is the reconstructed image. DETAILED DESCRIPTION
[0035] In order to more clearly express the purpose, technical solutions and advantages of the present invention, the specific implementation scheme of the present invention is described in detail below.
[0036] The present invention is a method for detecting the authenticity of an OCT fingerprint section image based on reconstruction difference. A fully convolutional neural network model is constructed, which includes three parts: an encoder, a generator, and a feature extractor, wherein the encoder and the generator are used to reconstruct the image. The reconstructed image shows a small reconstruction difference with the positive sample, but a large difference is shown when facing a negative sample. Considering that the difference is not accurate when directly reflected at the pixel level, and the feature encoding coupling of the encoder is relatively high, a feature extractor is set, in which channel attention and spatial attention modules are added to extract feature representations with more semantic information of the image. The cosine similarity is used to evaluate the feature similarity of the original input image and the reconstructed image after passing through the feature extractor. If the calculated result is lower than the set threshold, it can be judged as a forged image.
[0037] The OCT fingerprint section image authenticity detection method based on reconstruction difference of the present invention comprises the following steps:
[0038] Step S1, construct a fully convolutional neural network model. The main body of the model consists of three parts: an encoder, a generator, and a feature extractor, as shown in Figure 1. The encoder first obtains the feature map of the data distribution of the input image in the latent space, and then uses the generator to reconstruct an image similar to the input image from the obtained data distribution. Since the noise in the input image will increase the difference in pixels between the two images after reconstruction, it is inaccurate to use only pixel differences as a measure of true and false. The features extracted by the neural network can better reflect the main semantic information, which solves this problem to a certain extent. Since similarity evaluation is required in the feature space, the feature map information finally output by the decoder still has a high degree of feature coupling, retaining the background information in the original image, and it is difficult to directly use it as a feature representation of the image for subsequent feature comparison. Therefore, an additional feature extraction module is added to obtain a feature representation with more semantic information for the input image. This part uses RestNet as the basic structure. In order to more accurately locate the main area in the image and reduce background content interference, a channel attention module and a spatial attention module are added.
[0039] 1) Encoder.
[0040] First, use 8 convolution kernels of size f=3*3, set the step size s=1, and fill padding=1 on all sides for convolution operation, maintain the image size, and expand the number of input image channels to 8 channels.
[0041] Then, 5 layers of downsampling convolution layers are used. Each layer sets the convolution kernel size f=3*3, the step size s=2, and the padding=1. Instance Normalization is used for standardization. After each convolution layer, the image size is halved. Usually, the number of output channels, that is, the number of feature maps, is set to twice the number of original channels. Assume that the input size to the first downsampling convolution layer is BatchSize*Channel*Width*Height, where BatchSize is the number of each training batch, Channel is the number of channels, Width is the image width, and Height is the image height. A total of 5 layers of downsampling convolution layers are used in the model. Therefore, at the last layer, the output size should be BatchSize*(Channel*32)*Width / 32*Height / 32, that is, Channel*32 feature maps with a size of 32 times the original size are obtained.
[0042] 2) Generator.
[0043] This part consists of 5 upsampling layers, of which the upsampling layer consists of two parts and includes two processes: use the UpSample function to upsample and double the size of the feature map. Then use a convolution kernel of size f=3*3, set the step size s=1, and padding=1 around to perform the convolution operation, adjust the number of output channels, and usually set the number of output channels to half. The entire upsampling and image restoration process is similar to the reverse process of the encoder. The input of the first upsampling layer comes from the output of the last layer in the main branch of the encoder. Each time it passes, the number of channels (feature maps) is halved, and the size of the feature map is doubled. After 5 similar processes, the output size can be restored to the input image size, but the number of channels is 8 at this time, and it is still necessary to use convolution again to adjust the number of channels to 1, that is, to obtain the restored image, thereby realizing the restoration of the image from the extracted features.
[0044] 3) Feature extractor.
[0045] Using the ResNet network structure, the present invention adds channel and spatial attention mechanisms to it. The channel attention is generated by performing global maximum pooling and global average pooling on the feature map in the spatial dimension to obtain two C*1*1 vectors, then adding them, normalizing them with the Sigmoid activation function, and obtaining the final channel weight matrix. The spatial attention is generated by performing maximum pooling and average pooling on the feature map in the channel dimension to obtain two feature maps of size 1*W*H. A 7x7 convolution is used to keep the size of the feature map unchanged, and the feature map is fused into a feature map, and Sigmoid normalization is performed to obtain the final spatial weight matrix. After each layer of convolution, the feature map needs to be weighted with the corresponding channel and spatial weight matrices in turn.
[0046] Step S2: Prepare training data and test data. Collect B-scan (cross-section) fingerprint images collected by the OCT system, including 20 groups of real finger B-scan images and 10 groups of phantom B-scan images, each group of images has 400 images, from different individuals and different imitation materials. In addition, it is necessary to collect 10 images of the OCT system when the object to be tested is not placed, that is, images with only background. However, due to the poor quality of some original collected images, especially too much useless information on the left and right sides of the image, it is necessary to enhance the image before training. The specific process is: perform image cropping operation on the B-scan image of the original size of 1800*500, cut off 200 pixels on the left and right of the original image, and obtain a 1400*500 B-scan image, and then adjust the image size, use the bicubic interpolation method, and scale the cropped image size to the required size. In the experiment, it is scaled to 256*256 and converted to a grayscale image. After the preprocessing method, only 10 groups of real finger B-scan images were randomly selected from the 20 groups of real finger B-scan images as training data. Another 10 groups of real finger B-scan images and 10 groups of phantom B-scan images were selected as test data. For the 10 images containing only the background, data enhancement was performed to expand the number to 100 and save them for subsequent operations. The specific methods of data enhancement include: random cropping and then resizing to the original size, random Gaussian blurring, and random flipping.
[0047] Step S3: training the network model. The overall training process can be seen in the attached figure. Figure 2 . Select the divided training images as input data. Each time the data is loaded, the original image data is backed up and then randomly blocked with black blocks of random size. The image data obtained after the blocking operation is recorded as x′. It passes through the encoder E(*) and the generator G(*) to obtain the corresponding reconstructed image, recorded as G(E(x′)). Calculate the difference between the reconstructed image and the original unblocked input image in terms of pixels. The expected difference value is as small as possible so that the distribution of the generated image is as close as possible to the original input image. Use the L1 Loss mean absolute error, recorded as the reconstruction error L recon , calculated as follows:
[0048] L recon =||G(E(x′))-x|| 1 (1)
[0049] Where x represents the data distribution of the original input image, and G(E(x′)) represents the data distribution of the image reconstructed and restored after the network model. This loss function is only applied to the encoder and generator parts to improve the image reconstruction quality.
[0050] In order to alleviate the overfitting problem that may exist in the later stage of feature extractor training and improve the robustness of the model, a simple data augmentation operation can be used to amplify the data. It is necessary to vertically flip x and G(E(x)) to obtain the corresponding enhanced image data x^ and G(E(x))^. A total of 4 sets of data, including unenhanced and enhanced data, are input into the feature extractor. The obtained feature vector is taken as the positive feature vector, denoted as z pos At the same time, the same number of enhanced background image data prepared in step S2 is randomly selected and sent to the feature extractor. The feature vector obtained in this part is used as a negative feature vector, denoted as z neg . Start with z pos Select a positive eigenvector as the anchor point, denoted by z o , and then paired with another feature vector in the same batch. Among these combinations, the pairing combination of the anchor point and the positive feature vector is called a positive data pair, and the pairing combination with the negative feature vector is called a negative data pair. Assuming that the total number of feature vectors is M, after the above combination operation, 3 groups of positive data pairs, M-4 groups of negative data pairs, and a total of M-1 groups can be obtained. Then select the remaining positive feature vectors in turn and repeat the above operation.
[0051] The goal is to expect positive data pairs to have high similarity, while negative data pairs to have low similarity. The similarity of two vectors in a data pair is reflected by the cosine similarity calculation. The closer the value is to 1, the more similar the two vectors are, as shown in the following formula:
[0052]
[0053] Among them, S(a, b) is represented by the vector z a With vector z b Cosine similarity of data pairs, * T represents vector transpose, ||*|| represents the modulus of the vector, and γ is the scale parameter used to adjust the original [-1, 1] range of cosine similarity.
[0054] After determining the similarity measurement standard, set the contrast loss function L con , this loss function is similar to the softmax-cross entropy loss function in definition. In the process of optimizing the loss function, the proportion of positive data pairs similarity is gradually increased, so as to achieve the learning goal of the feature extractor part: maximizing the similarity of positive data pairs and minimizing the similarity of negative data pairs. First, calculate the proportion of positive data pairs composed of one anchor point in all combinations containing the anchor point. The goal is to expect the proportion to be as large as possible, so the loss function needs to take a negative sign again, as shown in the following formula:
[0055]
[0056] Among them, Lcon_anchor_n represents the average loss value of the positive data pair with the nth positive eigenvector as the anchor point, and M is the average loss value of the positive data pair with the anchor point z o_n The total number of positive data pairs, S(z o_n , z pos_i ) represents the i-th anchor point z o_n The cosine similarity of the positive data pair, N is the anchor point z o_n The total number of negative data pairs, S(z o_n , z neg_j ) represents the jth containing anchor point z o_n The cosine similarity of the negative data pairs.
[0057] Next, the loss values of the remaining anchor point combinations are calculated, and the above calculations are performed in sequence. Finally, the loss values obtained by all anchor point combinations are summed and averaged to obtain the final contrast loss L of the feature extractor part. con .
[0058]
[0059] Among them, N is the total number of anchor points set, and this loss function is only applied to the feature extractor part.
[0060] After setting the loss function, the constructed network model is trained for multiple rounds. The model weight parameters are updated and optimized through back propagation until the loss function tends to converge, at which time the training can be stopped.
[0061] Step S4: Test the network model. The overall test process can be seen in the attached figure. Figure 3 The pre-divided test data is selected as the input image, and the test is also divided into batches. Each time, 20 image data are randomly selected from them, that is, the size of each batch is 20*1*256*256, denoted as x, for testing. The testing process is similar to the training process in step S3. x passes through the encoder E(*) and the generator G(*) to obtain the corresponding reconstructed image, denoted as G(E(x)) (part of the reconstructed images during the experiment can be seen in the accompanying drawings Figure 4 and Figure 5), and is input into the feature extractor. The cosine similarity is also used to calculate the feature vector z corresponding to x and G(E(x)). 1 、z 2 The similarity is calculated and saved. Generally speaking, the similarity of positive samples is generally high, and that of negative samples is generally low. Then, the ROC curve is drawn based on the cosine similarity calculation results of all test data, and the appropriate threshold is set based on the accuracy, false detection rate, and missed detection rate. It is set that as long as the cosine similarity calculation is higher than the threshold, it can be identified as a real finger image, otherwise it is identified as a fake finger image.
[0062] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. An OCT fingerprint section image authenticity detection method based on reconstruction difference, characterized in that: The following steps are involved: Step S1, constructing a fully convolutional neural network model, the main body of the fully convolutional neural network model includes an encoder, a generator, and a feature extractor; the encoder obtains a feature map of the data distribution of the input image in the latent space; the generator reconstructs an image similar to the input image from the obtained data distribution; the feature extractor uses a ResNet network structure, in order to more accurately locate the main area in the image and reduce background content interference, a channel attention module and a spatial attention module are added to strengthen the acquisition of more semantic feature representation of the input image; Step S2, collecting images acquired by the OCT system, using B-scan images from real fingers of different individuals as positive sample images, and B-scan images from phantoms made of different imitation materials as negative sample images, and also collecting 10 images of the OCT system with only the background when no object to be tested is placed; then preprocessing these images, after the preprocessing is completed, randomly selecting 70% of the positive sample images from the positive sample images as training data; selecting another 30% of the positive sample images and negative sample images, and the number is balanced as test data; Step S3, training the network model; selecting the divided training images as input data, setting the loss function L recon , used to optimize the encoder and generator to improve the image reconstruction quality; set the contrast loss L con , used to optimize the feature extractor; conduct multiple rounds of training on the constructed network model, and update and optimize the model weight parameters through back propagation until the loss function tends to converge, and then stop the training; The loss function L is set recon , specifically including: The input data consists of two parts: the original input image data and the randomly occluded image data. The randomly occluded image data is obtained by occluding the corresponding image at a random position with a black block of random size each time the data is loaded. The occluded image is selected as the training data and sent to the encoder and generator. The original input image data is used as a measurement indicator, and the reconstructed output image needs to be compared with it, that is, the difference between the reconstructed image and the original unoccluded image at the pixel point is calculated. The expected difference value is as small as possible, so that the distribution of the generated image is as close to the original input image as possible. The L1 Loss mean absolute error is used, which is recorded as the reconstruction error L recon , calculated as follows: L recon =||G(E(x))-x||1 (1) Where x represents the data distribution of the original input image, and G(E(x)) represents the data distribution of the image reconstructed and restored after the network model. The setting contrast loss L con , the specific process is: Vertically flip the input and reconstructed images x and G(E(x)) to obtain the corresponding enhanced image data x^ 、 G(E(x))^, a total of 4 sets of data, including unenhanced and enhanced data, are input into the feature extractor, and the obtained feature vector is taken as the positive feature vector, denoted as z pos At the same time, the same number of enhanced background image data prepared in step S2 is randomly selected and sent to the feature extractor. The feature vector obtained in this part is used as a negative feature vector, denoted as z neg ; Start with z pos Select a positive eigenvector as the anchor point, denoted by z o , and then paired with another feature vector in the same batch. Among these combinations, the pairing combination of the anchor point and the positive feature vector is called a positive data pair, and the pairing combination of the anchor point and the negative feature vector is called a negative data pair. Then select the remaining positive feature vectors in turn and repeat the above operation. The similarity between the two vectors in the data pair is reflected by the cosine similarity calculation. The closer the value is to 1, the more similar the two vectors are, as shown in the following formula: Among them, S(a,b) is represented by the vector z a With vector z b Cosine similarity of data pairs, * T represents vector transposition, ||*|| represents the modulus of the vector, and γ is the scale parameter, which adjusts the original [-1,1] range of cosine similarity; After determining the similarity measurement standard, set the contrast loss function L con , this loss function is similar to the softmax-cross entropy loss function in definition. In the process of optimizing the loss function, the proportion of positive data pairs similarity is gradually increased, so as to achieve the learning goal of the feature extractor part: maximize the similarity of positive data pairs and minimize the similarity of negative data pairs; first calculate the proportion of positive data pairs composed of one anchor point in all combinations containing the anchor point. The target expectation is that the larger the proportion, the better, so the loss function needs to take a negative sign again, as shown in the following formula: Among them, L con_anchor_n represents the average loss value of the positive data pair with the nth positive eigenvector as the anchor point, and M is the average loss value of the positive data pair with the anchor point z o_n The total number of positive data pairs, S(z o_n ,z pos_i ) represents the i-th anchor point z o_n The cosine similarity of the positive data pair, N is the anchor point z o_n The total number of negative data pairs, S(z o_n ,z neg_j ) represents the jth containing anchor point z o_n The cosine similarity of the negative data pairs; Next, the loss values of the remaining anchor point combinations are calculated, and the above calculations are performed in sequence. Finally, the loss values obtained by all anchor point combinations are summed and averaged to obtain the final contrast loss L of the feature extractor part. con ; Among them, N is the total number of anchor points, and this loss function is only applied to the feature extractor part; Step S4, testing the network model; applying the trained network model, selecting test data as input model for testing, setting thresholds based on comprehensive accuracy, false detection rate, and missed detection rate, and performing authenticity discrimination on the input image according to the set thresholds in subsequent actual applications.
2. The method for detecting authenticity of OCT fingerprint section images based on reconstruction difference according to claim 1, characterized in that: The encoder in the network model of step S1 specifically includes: It includes 5 downsampling convolution layers, and each layer sets the convolution kernel size f=3*3, the step size s=2, and the padding=1. After each convolution operation, the image size is halved, and the number of output channels, that is, the number of feature maps, is the number of convolution kernels used in this layer, thereby achieving downsampling and dimensionality reduction.
3. The method for detecting authenticity of OCT fingerprint section images based on reconstruction difference according to claim 1, characterized in that: The generator in the network model in step S1 specifically includes: It includes 5 upsampling layers, of which the upsampling layer consists of two parts and includes two processes: using the Upsample function for upsampling to double the size of the feature map; then using a convolution kernel of size f=3*3, setting the step size s=1, and padding=1 on all sides for convolution operation, maintaining the size of the feature map, adjusting the number of output channels, and setting the number of output channels to be halved.
4. The method for detecting authenticity of OCT fingerprint section images based on reconstruction difference according to claim 1, characterized in that: The image preprocessing described in step S2 specifically includes: The B-scan image with an original size of 1800*500 was cropped by cropping 200 pixels on the left and right of the original image to obtain a 1400*500 B-scan image. The image size was then adjusted and the bicubic interpolation method was used to scale the cropped image to the required size. In the experiment, it was scaled to 256*256 and converted to a grayscale image. For 10 images containing only the background, data enhancement was performed to expand the number to 100 and save them for subsequent operations. The specific methods of data enhancement include: random cropping and then resizing to the original size, random Gaussian blurring, and random flipping.
5. The method for detecting authenticity of OCT fingerprint section images based on reconstruction difference according to claim 1, characterized in that: The authenticity determination criteria in step S4 specifically include: 1) In the forward propagation process, the test image and the reconstructed image x 1、 x2 is input into the feature extractor to obtain the feature vector z1z2; 2) Use cosine similarity to calculate the similarity of feature vectors z1 z2; 3) Draw the ROC curve based on the cosine similarity calculation results of all test data, and set the threshold based on the accuracy, false positive rate, and missed positive rate; 4) As long as the cosine similarity calculation is higher than the threshold, it is determined to be a real finger image, otherwise it is determined to be a fake finger image.
Citation Information
Patent Citations
Visual SLAM test method based on convolutional neural network
CN110555881A
High-voltage line insulator defect detection method based on generative adversarial network
CN112184654A