METHOD FOR AUTHENTICATING MATERIAL SUBJECTS BY GLASS PATTERN
Convolutional neural networks sensitive to Glass patterns address the challenge of high computational demands in authentication by ensuring efficient and accurate unit authentication with flexible image acquisition, even under varying conditions.
Patent Information
- Application Number
- FR2024008783
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-13
AI Technical Summary
Existing methods for unit authentication of physical objects require significant computing power and long computation times, which is not feasible with limited resources and energy consumption, and lack flexibility in image acquisition conditions.
A method using convolutional artificial neural networks sensitive to Glass patterns in images with random textures, allowing authentication by comparing reference and candidate images with residual geometric transformations, even under varying acquisition conditions.
Achieves satisfactory computation times and lower energy consumption while maintaining high authentication accuracy, enabling reliable unit authentication on portable devices with limited resources.
Smart Images

Figure 00000032_0000 
Figure 00000033_0000 
Figure 00000033_0001
Abstract
Description
Title of the invention: METHOD FOR AUTHENTICATING MATERIAL SUBJECTS BY GLASS PATTERN
[0001] The present invention relates to the technical field of unitary authentication of material subjects with electronic means.
[0002] In the above-mentioned field, it is known, notably from US patent 4,423,415, to identify physical objects by extracting a signature from a so-called authentication region comprising an essentially random three-dimensional intrinsic microstructure. Initially, a first extraction, of a so-called reference signature of an authentic object, is generally performed using electronic computing means such as a computer, smartphone, or tablet, which implement complex pattern recognition algorithms. The reference signature is then recorded. Subsequently, a second extraction, of a so-called candidate signature of a candidate object, is performed using electronic computing means similar to those used in the first step.The reference and candidate signatures are then compared using similarity or correlation comparison algorithms to determine their level of proximity and, if applicable, to deduce that the candidate subject is indeed the authentic subject.
[0003] US patent 10,990,845 proposed a method for unit authentication of physical objects by calculating similarity vectors between, in particular, reference and candidate images of authentication regions of authentic and candidate physical objects. The methods and algorithms implemented in this patent require significant computing power for real-time implementation or long computation times with less power if high execution speed is not required. However, this constraint of significant computing power or long computation times proves to be a disadvantage when it is necessary to obtain computation times as short as possible while having limited computing power combined with low energy consumption.
[0004] It therefore arose the need for alternative solutions which provide a solution to this problem and allow for reliable unit authentication of material subjects while using limited computing resources and contained energy consumption and which offer greater flexibility with regard to the conditions of acquisition of reference and candidate images than prior art processes implementing signature extraction.
[0005] To achieve this objective, the invention relates to a method for authenticating a material subject consisting of comparing, on the one hand, at least one so-called reference image of at least one authentication region of an authentic subject, the reference image comprising at least one texture with a random component, and on the other hand, a so-called candidate image of at least one authentication region of a candidate subject, the candidate image comprising at least one texture with a random component. According to the invention, the method implements at least one artificial neural network configured to be sensitive to Glass patterns in the context of images with a texture with a random component. Thus, according to the invention, the determination of authenticity is based on the result provided by the implemented neural network(s).This result can, for example, be a numerical value, in which case a threshold can be used above which the authentication result is positive and below which there is no authentication.
[0006] In a preferred embodiment, the invention implements at least one convolutional type artificial neural network.
[0007] For the purposes of this invention, a convolutional artificial neural network, also called a convolutional artificial neural network, is a type of neural network comprising at least one, and preferably several, convolutional layers, and after the last convolutional layer, at least one layer of artificial neurons of another type, for example, at least one fully connected or densely connected layer of artificial neurons. Hereafter, the terms "convolutional artificial neural network," "convolutional neural network," and "convolutional neural network" shall be used interchangeably as synonyms. Similarly, within the scope of this invention, the terms "neural network" and "artificial neural network" are synonymous, and the term "neuron(s)" shall be understood to mean "artificial neuron(s)" unless otherwise specified.
[0008] The implementation of convolutional artificial neural networks makes it possible to obtain satisfactory computation times with reasonable computing resources, such as those available in portable devices like smartphones or tablets, while exhibiting lower energy consumption than that required by prior art computational methods. Furthermore, the implementation of artificial neural networks allows for a certain flexibility in the acquisition conditions of reference and candidate images.
[0009] For the purposes of the invention, it should be understood that a convolutional artificial neural network configured to be sensitive to Glass patterns in the context of images with random texture components is an artificial neural network that has been trained at least with images whose superposition reveals or is likely to reveal a pattern analogous to a Glass pattern, as has been described in patent EP 3 380 987 in the name of the applicant, to which reference should be made. Within the framework of the invention, the artificial neural network(s) implemented are configured to react to images which, in combination, produce or are likely to produce said Glass pattern, or to an image resulting from the superposition of two images and producing said Glass pattern.
[0010] Thus, the present invention takes advantage of the applicant's demonstration that patterns similar to those obtained by Léon GLASS in articles in the journal NATURE vol.223 of August 9, 1969 pages 578 to 580 and Nature vol. 246 of December 7, 1973 pages 360 to 362, can appear by superimposing two images comprising respectively natural textures or random components resulting from the acquisition or even the photography at an appropriate magnification or enlargement of the same intrinsic random multi-scale three-dimensional material structure of the same subject.The applicant has demonstrated that these Glass-type patterns appear only when continuous random textures originating from the same material structure and essentially residual geometric transformations of each other are superimposed. In practice, they do not appear when the continuous random textures are not sufficiently correlated or do not result from the acquisition of the same material structure corresponding to a subject's recognition region. Therefore, within the scope of the invention, Glass patterns enable unitary authentication, also known as unitary recognition. Conversely, if a Glass pattern is not observed, it is not possible to definitively conclude that the data is not authentic.Furthermore, the applicant highlighted the fact that, in the case of physical subject authentication, if said physical subject exhibits sufficient stability over time, images taken at different times, even if separated by several days, months, or years, can, through their superposition, generate such Glass patterns. Moreover, according to the invention, the authentic subject can undergo modifications after the authentication image is recorded while remaining authenticable, provided that a portion of the recognition region has not been significantly affected by these modifications, whether intentional or not.Furthermore, the applicant has demonstrated that the observation of Glass patterns is possible by superimposing two images of the same recognition region possessing a texture with a continuous random component, without the addition, extraction, or generation of discrete elements or discrete patterns as advocated in the prior art before patent EP 3,380,987 in the applicant's name. Such a Glass pattern appears only in the case of an authentic subject and only if there exists a non-zero slight geometric transformation, referred to in the sense of the invention as a residual geometric transformation, between the candidate subject and the authentication image, or between the verification and authentication images acquired under the given conditions. This... The properties of Glass patterns offer great robustness to the method according to the invention, insofar as the acquisition conditions of the verification image do not need to be strictly identical to the acquisition conditions of the authentication image. Thus, the resolutions of the authentication and verification images may differ. In the context of this application, an authentication image is also referred to as a reference image, while a verification image is also referred to as a candidate image.
[0011] Within the framework of the invention, an authentication region, also called a recognition region, is a region that exhibits an intrinsic and random microstructure in that it results from the very nature of the authentication region of the material subject. In a preferred embodiment of the invention, each material subject used belongs to the subject families comprising at least one authentication region having an essentially random intrinsic structure that is not easily reproducible, that is to say, whose reproduction is difficult or even impossible in that it results in particular from a process that is not predictable at the observational scale.Such an authentication region with an essentially random, non-easily reproducible intrinsic continuous medium structure corresponds to non-clonal physical functions, also known as "Physical Unclonable Functions" (PUFs), as defined in particular by the English publication Encyclopedia of Cryptography and Security, 01 / 2011 edition, pages 929 to 934, in the article by Jorge Guajardo. Preferably, the authentication region of a material subject conforming to the invention corresponds to an intrinsic non-clonal physical function designated in English as "Intrinsic PUFs" in the aforementioned article.The applicant takes advantage of the fact that the random nature of the authentication region's microstructure is inherent or intrinsic to the very nature of the subject, resulting from its mode of formation, development, or growth. Therefore, it is unnecessary to add a specific structure to the authentication region, such as an impression or engraving. Regarding the recognition region, according to the invention, its image possesses a texture that an observer with average visual acuity can perceive either with the naked eye or via optically and / or digitally enhanced zoom. Thus, small structures, or structures perceived as small at the observation magnification, are seen as images containing a texture.
[0012] In the context of the invention, the term texture refers to what is visible or observable in an image, while the term structure or microstructure refers to the material subject itself. The texture of interest in the recognition region is described as a texture with a random component in that it includes at least some randomness or irregularity. Thus, a texture or The microtexture of the recognition region, known as a random-component texture image, corresponds to an image of the structure or microstructure of the recognition region. Imaging such a recognition region under similar observation conditions from neighboring viewpoints yields images, each containing a random-component texture that is a noisy reflection of its material structure. Such a random-component texture inherits its unpredictability and independence from a random-component texture originating from a completely different recognition region, due to the degree of randomness in the formation of their material structures.
[0013] For the purposes of this invention, the term "image," whether candidate, reference, or training, means any type of image in the general sense of the term, and not solely an image comparable to a photograph. In other words, an image is not limited to an optical image resulting from the stimulation of the recognition region by visible light, but can instead be obtained by any type of physical stimulation, including but not limited to: ultrasound, far-infrared, terahertz, X-rays or gamma rays, X-ray or laser tomography, X-ray radiography, and magnetic resonance imaging. Thus, for the purposes of this invention, an image is, for example, the recording of the result of stimulation by any means whatsoever of a natural scene or a material subject. This recording can then be described as a natural image.This recording may be one-dimensional, corresponding, for example, to the recording of the variation over time or along a line of a single signal, or to the recording of the values from a line of sensors. This recording may also be two-dimensional, as is the case with a photograph, which can be recorded in halftones, grayscale, or color. For the purposes of this invention, an "image" can therefore be a 1D signal, a 2D or 3D image, in grayscale or color, or an nD (n-dimensional) signal, for example, hyperspectral or RGB-D. For the purposes of this invention, an "image" can be the result of a single acquisition or extracted from a stream. Thus, within the scope of this invention, an "image" can be extracted from a video stream. In one embodiment, the image can be saved digitally.Furthermore, an image implemented in the process according to the invention may be a pre-segmented image, particularly when it comprises several material subjects in the same scene or a repetition of the same motif. Of course, an image is not necessarily natural and may be synthetic, that is to say, it may be generated by a computer process with or without the assistance of a human operator. Within the scope of the invention, a natural image and a synthetic image are considered. They have in common that they are in the same digital or analog recording format for processing within the same process. Optical and / or digital image enhancement preprocessing can also be applied to improve the signal-to-noise ratio, for example.
[0014] According to a first embodiment of the invention, the artificial neural network is provided with an image resulting from the superposition of the candidate image and the reference image, the artificial neural network having been trained using images with random textures exhibiting a Glass pattern. As part of its training, the neural network according to this first embodiment will also be trained with superimposed images of random textures not exhibiting a Glass pattern.
[0015] Image superposition refers to any combination of all or part of at least two images to obtain a new image in which all or part of the original images contribute. According to a preferred feature of the invention, the candidate authentication, reference, and verification images do not undergo, for the purpose of superposition, any transformation during the verification phase other than operations to enhance or modify contrast, brightness, halftone transformation, color space changes such as conversion to grayscale or black and white, operations to modify saturation in certain hues, level inversion, or relative opacity modifications via the alpha channel.Thus, according to this preferred characteristic, images generally undergo so-called enhancement transformations that do not affect the ability to visually recognize the nature of the subject. Preferably, the applied transformations do not distort the images, in particular the continuous random textures they contain.
[0016] In this context, superimposition consists in particular of: • superimpose one image on another by aligning the two images and stacking them on top of each other • Calculate the difference between two images by subtracting the pixel values of one image from the other / Weighted average: Calculate a weighted average of the two images • Blend two images using a blending function. This method combines the color channels of the two images to create a blending effect. • Perform a layering technique using "alpha-blending": This method involves using an alpha channel to control the transparency of the layered image. • perform a superposition with offset, namely offsetting the superimposed image relative to the base image, • perform a superimposition with a scaling modification, namely resize the superimposed image to compare areas of interest of different sizes, • perform a superposition with variable opacity, namely adjust the opacity of the superimposed image to visually compare the images according to the desired opacity, • Perform a mask overlay, that is, create a mask to specify the areas to be overlaid in each image. It should be noted that a Glass pattern may appear when viewing the overlay of more than two images. Therefore, within the scope of this invention, the term "overlay" refers to the superposition of at least two images.
[0017] According to a variant of the first embodiment, prior to their superposition, the candidate and reference images are recalibrated with respect to each other.
[0018] In the context of this variant and according to a feature of the invention, after registration and before their superposition the candidate and reference images are subjected to a residual transformation with respect to each other.
[0019] For the purposes of this invention, a residual transformation or residual geometric transformation is a non-zero, lightweight geometric transformation to be applied locally to the chosen image or images, whether rigid or not, linear or non-linear, at at least one fixed or quasi-fixed point. Among the applicable geometric transformations, it is thus possible to implement the transformations described by Leon Glass in his 1973 and 2002 articles cited above. A quasi-fixed point is defined as a point that, after residual geometric transformation, undergoes a displacement of small amplitude compared to the maximum displacement caused by the residual geometric transformation. In the theoretical case of a perfect superposition of strictly identical elements / images, a Glass pattern does not appear, even in the presence of an authentic subject. This underscores the necessity of this residual geometric transformation and the general advantage of implementing a relative movement or displacement, or even a deformation induced by a difference in shooting angle or viewpoint between the authentication and verification images.
[0020] According to a second embodiment of the invention, the candidate image and the reference image are provided to the artificial neural network without superimposition of these, the artificial neural network having been trained by means of pairs. of non-overlapping images with random textures, the superposition of which produces a Glass pattern, referred to as training pairs. In this second embodiment, and according to a variant of the invention, training is performed using, on the one hand, images known to generate a Glass pattern when superimposed, without any special precautions regarding their presentation to the artificial neural network to be trained, and, on the other hand, images known not to generate Glass patterns as defined in the invention when superimposed. By "no special precautions," it is understood that the images in a pair have not undergone any registration, either during the training or processing phases.
[0021] By proceeding in this way, an artificial neural network configured to be sensitive to the Glass pattern was obtained, achieving an identification rate of approximately 80% during operation, without it being possible to substantially improve this authentication rate by increasing the size of the training datasets. By "images whose superposition is known to be likely to generate Glass patterns," it should be understood that these are images from the superposition of which a human operator will certainly observe a Glass pattern, possibly after registration operations have been performed, for example, manually.
[0022] Such an identification rate of 80% may be satisfactory in certain applications, particularly when it is possible to have a second check performed by an operator implementing the invention as described in patent EP 3#380#987 on behalf of the applicant. However, this rate is insufficient when it is necessary to process a large number of physical subjects, or their images, within a limited and short timeframe, thus restricting the possibility of human intervention.
[0023] The need therefore arose to improve the authentication or recognition rate. To this end, a variant of the second embodiment of the method according to the invention proposes to register the images of each training pair with respect to each other prior to their submission to the artificial neural network during the training phase. It was found that, quite surprisingly, implementing registration and then applying a residual transformation prior to presenting the images of the training pairs yielded better results than when no precautions were taken. By proceeding in this way, it was possible to achieve authentication rates of around 95% or even higher during the operational phase.
[0024] According to a characteristic of this variant, after registration and before supplying the artificial neural network, the images of each training pair are subjected to a residual transformation with respect to each other.
[0025] According to a characteristic of the second embodiment of the invention, prior to their provision to the artificial neural network, the candidate and reference images are calibrated against each other.
[0026] Preferably but not exclusively, within the framework of this feature, subsequent to registration and prior to their superposition, the candidate and reference images are subjected to a residual transformation with respect to each other.
[0027] According to a feature of the second embodiment of the invention, the two images to be processed, either the training pair or the reference image and the candidate image, are assembled by juxtaposing them without overlap to form a single image that will be provided to the artificial neural network. Preferably, this assembly is performed after any necessary registration, possibly followed by a residual transformation. In the context of the invention, assembly by juxtaposition without overlap corresponds to concatenation.
[0028] According to a third embodiment of the invention, two twin artificial neural networks operating in parallel are implemented, the reference image being provided to a first network and the candidate image being provided to the other network, the two networks having been trained by means of pairs of non-superimposed random component textured images whose superposition reveals a Glass pattern, called training pairs, the two images of the same pair being provided each to a separate artificial neural network.
[0029] According to a characteristic of this third embodiment, prior to their provision to the artificial neural networks, the images of each training pair are recalibrated with respect to each other.
[0030] According to a variant of this feature, after registration and before being supplied to the artificial neural networks, the images of each training pair are subjected to a residual transformation with respect to each other.
[0031] According to another feature of this third embodiment, the reference and candidate images are calibrated against each other before being supplied to the artificial neural networks.
[0032] According to a variant of this feature, after registration and before being supplied to neural networks, the candidate and reference images are subjected to a residual transformation with respect to each other.
[0033] According to a characteristic of the method according to the invention, each candidate image is derived from a video stream. This characteristic allows, in certain contexts of implementation The aim is to avoid residual registration and transformation operations applied to candidate and reference images prior to their processing by each neural network. This is particularly relevant when the video stream originates from an acquisition of a physical subject to be authenticated, performed by an operator who carries out crude registration during the acquisition process.
[0034] According to an embodiment of the method according to the invention, each artificial neural network is adapted to process images or imagelets of given dimensions and when the images to be processed are larger, the images to be processed are cut into sub-images or imagelets of suitable size which are successively submitted directly to each artificial neural network.
[0035] In the context of this application and when referring to the processing performed by the invention, the term "image" is preferably used for the result of the acquisition, while the term "imagelet" is used for the object that is processed by each neural network. However, it should be noted that images and imagelets are of the same nature, differing only in size or dimensions, so that these terms may be used interchangeably depending on the context, without hindering the understanding of the invention by a person skilled in the art. Indeed, the terms "image(s)" and "imagelet(s)" are used to facilitate understanding of the description of the invention and should therefore not be interpreted as being restrictive.Similarly, in some cases, the sub-images may also be referred to as the block; this term, also used to facilitate the description of the invention, should not be interpreted in a restrictive manner.
[0036] According to a feature of this method of implementing the process according to the invention, the candidate and reference images are recalibrated prior to their cutting.
[0037] According to a preferred feature of the invention, the registration is carried out before the cutting while the residual transformation is applied after the cutting.
[0038] According to a preferred feature of the invention, at least one implemented artificial neural network is a convolutional neural network.
[0039] The invention also relates to a computer program product comprising code instructions for executing a method according to the invention for the one-time authentication of a physical subject from a reference image and a candidate image, when the program is executed on a computer. Such a computer program product can then be embedded or stored in a computer or the like. For the purposes of the invention, the term "computer" should be understood in a broad sense as including, in particular, a personal computer, a smartphone, a tablet, a virtual or physical server, or any computing unit and a process capable of implementing a computer program and adapted to the implementation of the invention.
[0040] The invention also relates to a means of storage readable on a computer equipment on which a computer program includes code instructions for the execution of a method according to the invention of authenticating a material subject from a reference image and a candidate image.
[0041] The invention also relates to a computer device comprising at least display means, image acquisition means, user information input means, information storage means communicating with computing and control means configured to implement the method according to the invention of authenticating a material subject from a referenced image of a candidate image.
[0042] The different modes of implementation, forms of embodiment, characteristics and variants of the invention can be implemented with each other in different combinations insofar as they are not mutually exclusive or incompatible with each other.
[0043] Various other features and variants of the invention will become apparent from the description below, made in relation to the figures in which: Figure 1 is a schematic representation of a first embodiment of the invention implementing a convolutional neural network. Figure [Fig. 2] illustrates an example of slicing the superposition of a reference image and a candidate image. Figure 3 illustrates another example of clipping and processing applied to the reference and candidate images used in Figure 2; Figure 4 shows image sets of training sets used for training convolutional neural networks implemented within the scope of the invention. Figure 5 is a schematic representation of a second embodiment of the invention. Figure 6 is a schematic representation of the operation of an image slicing module implemented in the second embodiment of the invention. Figure 7 is a schematic representation of a third embodiment of the invention implementing two artificial neural networks. Figure 8 is a schematic representation of a variant of the third embodiment of the invention implementing three artificial neural networks.
[0044] In the figures, the elements common to the different embodiments bear the same reference numerals. Furthermore, the different embodiments, presented in relation to the figures, correspond to non-limiting examples of possibilities for execution and implementation of the invention.
[0045] As previously stated, the invention implements at least one convolutional artificial neural network to ensure the authentication of physical objects from images of an authentication region of these objects, the images comprising at least one texture with a random component. In a preferred embodiment, the invention enables unitary authentication, that is, a reference image, also called an authentication image, corresponds to one and only one physical object.
[0046] Thus, unit authentication means unit recognition of a region of a material subject. This recognition can have a higher or lower probative value depending on the criticality of the use case considered and the measures implemented to increase this probative value, such as, but not limited to: the involvement or not of a trusted third party, the selection of highly sophisticated acquisition sensors in terms of resolution and illumination conditions for acquiring reference images and the use of the same sensors for acquiring candidate images, and complete control of the IT environment implemented, without this list being exhaustive or limiting.
[0047] According to the invention, the authentic subject can undergo modifications after the recording of the reference image, also called the authentication image in the context of patent EP 3#380#987, while remaining authenticable to the extent that a part of the authentication region has not been profoundly affected by these voluntary or involuntary modifications.
[0048] In order to be implemented on portable devices and / or to avoid requiring significant computing resources, which are also energy-intensive, the invention proposes implementing artificial neural networks that could be described as frugal in that they comprise a limited or even reduced number of layers. In a preferred embodiment, the invention implements convolutional neural networks comprising a limited number of convolutional neuron layers and fully or densely connected neuron layers.
[0049] Two examples of convolutional neural networks, also known by the abbreviation "CNN" (Convolutional Neural Network), that could be implemented by the invention include LetNet, LetNeT-5, and AlexNet. The French and English Wikipedia pages entitled "convolutional neural network" and "convolutional neural network," respectively, provide further examples of convolutional neural networks and explanations of their structure.
[0050] According to the preferred embodiment of the invention, one or more convolutional neural networks are used, each comprising a stack of processing layers, namely: • Convolutional layers (CONV) that process data from a receiving field, namely an image, • POOL pooling layers that allow information to be compressed by reducing the size of the intermediate image • correction layers often incorrectly called ReLU by reference to the rectified linear activation function, • completely or densely connected layers, FDC • LOSS layer, which can also be called a loss layer.
[0051] It should be noted that for some authors the LOSS layer is considered not to be part of the neural network; therefore, within the scope of the invention, a neural network does not necessarily have such a layer. Furthermore, the so-called LOSS layer is present only during the learning phase and not during the execution phase.
[0052] Thus, and as can be seen from [Fig. 1], an example of such a convolutional artificial neural network, designated collectively as CNN, is configured to process a rectangular image img of mxn pixels, preferably with m and n greater than or equal to 64, for example, 128x128 pixels, it being understood that m and n are not necessarily equal. The CNN neural network comprises a convolution processing block 1 followed by a processing block of fully or densely connected artificial neurons 2.
[0053] According to the illustrated example, the CNN network further comprises, between the convolution block 1 and the processing block 2 with completely or densely connected artificial neurons, a FLAT layer for processing the output of the convolution block 1 before supplying it to the block 2.
[0054] According to the illustrated example, the convolution block 1 comprises successively and in this order: - a first convolution layer 11, - a first correction layer 12, - a first layer of pooling 13, - a second convolution layer 14, - a second correction layer 15, - a second layer of pooling 16, - a third convolution layer 17 - a correction layer 18 which happens to be the last layer of convolution block 1.
[0055] The first convolution layer 11 is, in this case, a 2D convolution layer parameterized to process the entire image using 3x3 pixel tiles (kernels or filters) with a step of 1 pixel. The first correction layer 12 implements a Rectified Linear Unit (ReLU) activation function to process the result of the first convolution layer 11. The first pooling layer 13, as illustrated in the example, performs maximum pooling for 2D spatial data with a 2x2 window and a step (Stride) of 2, on the result of the processing performed by the first correction layer 12.
[0056] The second convolution layer 14 is, in this case, a 2D convolution layer parameterized to process the output of the first pooling layer 13 by 5x5 pixel tiles with a step of 2 pixels. The second correction layer 15 implements a Rectified Linear Unit (ReLU) activation function to process the result of the second convolution layer 14. The second pooling layer 16 ensures, according to the illustrated example, a maximum pooling operation for 2D spatial data with a 2x2 window and a step of 2 on the result from the second correction layer 15.
[0057] The third convolution layer 17 is, in this case, a 2D convolution layer parameterized to process the result of the second pooling layer 16 by 5x5 pixel tiles with a step of 2 pixels. The third correction layer 18 here implements a Rectified Linear Unit (ReLU) activation function to process the result of the third convolution layer 17. In the illustrated example, the third correction layer 18 is the last layer of the convolution block 1.
[0058] An example of code for defining convolution block 1 as described above in PyTorch is as follows: self.cnnl = nn.Sequential(nn.Conv2d(l, 128, kernel_size=3, stride=l), nn.ReLU(inplace=True), nn.MaxPool2d(2, stride=2), nn.Conv2d(128, 256, kernel_size=5, stride=2), nn.ReLU(inplace=True), nn.MaxPool2d(2, stride=2), nn.Conv2d(256, 512, kernel_size=5, stride=2), nn.ReLU(inplace=True), )
[0059] The CNN neural network includes, at the output of convolution block 1, a FLAT processing layer, having reference 19, which ensures the flattening in the form of a single-row vector or matrix the result of the processing from the third correction layer 18.
[0060] Downstream of the processing layer 19, the CNN neural network comprises the processing block 2, which is a fully or densely connected neural network. According to the illustrated example, block 2 comprises successively and in this order: - a first layer 20 of densely or fully connected neurons - a first correction layer 21, - a second layer 22 of densely or fully connected neurons, - a second correction layer 23, - a third layer 24 of densely or fully connected neurons, which, according to the illustrated example, is the last layer of block 2.
[0061] According to the illustrated example, the first layer 20 is a linear layer of artificial neurons receiving 512 inputs and delivering 1024 outputs. By linear layer, it is understood that a line of neurons, that is to say, a layer with a single thickness of neurons.
[0062] The first correction layer 21, of block 2, here implements an activation function of type ReLU for in English "Rectified Linear Unit" applied to each of the 1024 outputs of layer 20.
[0063] The second layer 22 is a linear layer of artificial neurons receiving 1024 inputs and delivering 256 outputs.
[0064] The second correction layer 23, of block 2, here implements an activation function of type ReLU for in English "Rectified Linear Unit" applied to each of the 256 outputs of layer 22.
[0065] Finally, the third and last layer 24 is a linear layer of artificial neurons receiving 256 inputs and delivering 1 output.
[0066] An example of code for the definition of block 2, for a fully or densely connected neural network, as described above in PyTorch language is as follows: self.fcl = nn.Sequential( nn.Linear(512, 1024), nn.ReLU(inplace=True), nn.Linear(1024, 256), nn.ReLU(inplace=True), nn.Linear(256, 1), )
[0067] According to the illustrated example, the CNN neural network finally includes an output layer 24 which, in this case, ensures a normalization of the value delivered by the third and last layer 24 of block 2, in the form of a floating-point real number, between 0 and 1.
[0068] It should be noted that the different values of the functions mentioned in the Pytorch code examples are designated by the generic term hyperparameters; these are data which are not automatically updated during the learning phases, also called training phases, as opposed to parameters which are, such as, for example, the weight of the connections between artificial neurons and the biases of these.
[0069] According to the illustrated example, an input block (INPUT) is implemented upstream of the CNN. This block performs various processing operations on each image or pair of images to be processed before they are supplied to the CNN neural network. For example, when the image to be processed (IMG) is larger than the image (img) that can be processed by the CNN neural network, the input block (INPUT) will perform a segmentation or decomposition of the image to be processed (IMG) into sub-images or imagelets (img) that will be directly supplied to the CNN neural network. In other words, all processing operations that can be performed by the input block (INPUT) will be carried out prior to this segmentation. In the context of this application, the segmentation performed by the segmentation module corresponds to segmenting an image into a multitude of smaller imagelets or into a multitude of smaller blocks.
[0070] The training, also called learning, of an artificial neural network takes place according to a number of phases well known to those skilled in the art, which are as follows: 1 / Initialization of weights: The weights of the connections between neurons are initialized with small random values. 2 / Presentation of training images: The training images are introduced into the network, starting with the input layer and propagating towards the output layer of the network. 3 / Output calculation: The network calculates the output for the given input, by applying non-linear activation functions in each neuron. 4 / Error calculation: The error between the actual output and the predicted output is calculated, usually using a loss function. 5 / Backpropagation of error: The error is propagated back through the network, through each layer, by adjusting the weights of the connections between neurons. 6 / Weight update: The weights are adjusted according to the error gradient, using an optimization method such as gradient descent. 7 / Repetition: Steps 2 to 6 are repeated for a large number of epochs (iterations) until the network achieves acceptable accuracy on the training images. 8 / Evaluation / Test: The trained network is evaluated on a set of test images to estimate its performance.
[0071] An "epoch" in the context of machine learning corresponds to a complete pass through the CNN of all the training (respectively, test) images or pairs of training images. In other words, an epoch is a complete iteration in which the model sees all the training (respectively, test) images or pairs of training images once, and the weights and biases are updated accordingly. For example, if the set of training (respectively, test) images or pairs of training images contains 1000 images or pairs of images, an epoch corresponds to the presentation of these 1000 images or pairs of images to the CNN, and the updating of the weights and biases after each image. It is important to note that the term "epoch" is often used interchangeably with the term "iteration," but they have slightly different meanings.An iteration can correspond to a presentation of a single image or pair of images in the context of the invention, while an epoch corresponds to a complete pass through the CNN of all the images or pairs of images.
[0072] The concept of convergence (training and testing phases) is essential for validating the learning process. Convergence on training images is used to adjust the hyperparameters of the network model and to determine whether the model has converged, while convergence on test images is used to evaluate the final performance of the model on unknown images. The convergence of a neural network is the process by which the network learns to represent the relationships between inputs and outputs, and where the weights and biases of the neurons are adjusted to minimize prediction error.
[0073] It is important to note that the convergence of a neural network is not always guaranteed, and that it is possible that the network may not converge towards an optimal solution.
[0074] Training convergence occurs when the model's loss (or error) on the training images or pairs of images decreases over the course of training iterations (epochs) and reaches a plateau. This means that the model has learned to represent the relationships between the inputs and outputs in the training images.
[0075] Convergence of the test occurs when model loss on the test images decreases over the course of training iterations (epochs) and reaches a plateau. Model loss refers to the images or pairs of images that were not recognized when they should have been. This means that the model generalizes well to new images that it did not see during training.
[0076] The objective is to achieve simultaneous convergence of the training and testing phases, which means that the model learns to represent the relationships between inputs and outputs in the training images and generalizes well to new images.
[0077] If convergence in the training phase is rapid, but convergence in the testing phase is slow or does not occur, this may indicate that the model is overfitted. That is, the model is too specialized for the training images and does not generalize well to new images.
[0078] Conversely, if convergence in the test phase is rapid, but convergence in the training phase is slow or does not occur, this may indicate that the model is underfitted. This means that the model has not sufficiently learned the relationships between the inputs and outputs in the training images.
[0079] In summary, the convergence of the training and testing phases is an important indicator of the performance of a CNN model, and it is essential to monitor these two metrics to adjust the hyperparameters and improve the performance of the implemented CNN model. "Full" learning includes both the training and testing phases; therefore, a neural network is only put into operation after undergoing both training and testing.
[0080] Once the CNN neural network has been trained, it is implemented to ensure, in a phase, called exploitation or execution, and in accordance with the invention, the authentication of a material subject by confronting, on the one hand, at least one so-called reference image of at least one authentication region of an authentic subject, the reference image comprising at least one texture with a random component, and on the other hand, a so-called candidate image of at least one authentication region of a candidate subject, the candidate image comprising at least one texture with a random component.
[0081] Thus, for each authentication operation, at least one image acquired from a recognition region of a candidate material subject, called the candidate image, will be used, and at least one reference image acquired from a recognition region of a previously recorded reference material subject, called the reference image.
[0082] EP patent 3,380,987 sets forth acquisition conditions for the reference and candidate images suitable for their superposition to produce a Glass pattern. It is particularly important to note from this patent that the Glass pattern appears only in the case of an authentic subject and that there exists a slight non-zero geometric transformation, referred to in the invention as a residual geometric transformation, between the verification (candidate) and authentication (reference) images acquired under the given conditions. In the case In theory, a perfect superposition of strictly identical elements / images does not result in the appearance of a Glass pattern, even in the presence of an authentic subject. This underscores the necessity of this residual geometric transformation and the general advantage of implementing a relative movement or displacement, or even a deformation induced by a difference in shooting angle or viewpoint between the authentication and verification images. This property of Glass patterns provides significant robustness to the method according to the invention, as it is not necessary for the acquisition conditions of the candidate verification image to be strictly identical to the acquisition conditions of the reference authentication image. Thus, for example, the resolutions of the reference authentication and candidate verification images can be different.
[0083] In the case of the method according to the invention, the reference and candidate images are compared in a process implementing at least one neural network such as, for example, described above and which has been trained as stated. By "compared" it should be understood that the result of the process depends on these two images and that it is absolutely necessary to have them in order to implement the invention, and that the reference and candidate images are two input data for the method according to the invention.
[0084] Note that large source, candidate, and reference images can be used, but they must then be cropped to fit the imagelets that the chosen CNN accepts as input. For example, if the candidate and reference images are 3840 x 2880 pixels and the CNN is configured to process 128 x 128 pixel imagelets, the candidate and reference images can be cropped into 660 imagelets. To this end, and as shown in [Fig. 1], an input block INPUT is implemented before the CNN. This block includes a cropping module ECH and successively provides each of the resulting cropped imagelets to the CNN.
[0085] Given the large number of processed images, unit recognition of the authentication region will be considered achieved if a given percentage of the candidate images have been recognized as implementing the Glass phenomenon on the two original images considered. In this regard, it should be noted that authentication can be concluded if the Glass phenomenon is identified for at least one presented candidate image. In order to perform the processing involved in this decision method, an OUTPUT block is implemented after the CNN network. This block collects the CNN network output for each of the images processed by the CNN and performs the processing necessary to make a decision regarding the identification or not of images showing a Glass phenomenon or likely to cause one to appear in case of superposition with or without registration.
[0086] According to one variant, the OUTPUT block can implement a statistical comparison with a similarity index between images, determine a given acceptance threshold, per pair of images or for a set of images located or not on the material subject, and issue an acceptance decision possibly associating it with a confidence index of this decision.
[0087] In a first embodiment of the invention, more particularly illustrated in [Fig. 1], the reference image IMG Ref and the candidate image IMG Cand are submitted to the CNN neural network after being superimposed on one another. For this purpose, the input block INPUT comprises a first module 30 which first performs registration of the reference and candidate images and then induces a residual transformation of one of these two images with respect to the other.
[0088] Registration refers in particular to coarse registration, that is, an approximate alignment of images with each other, based on general characteristics such as size, orientation, or overall position. This can be carried out, for example, during the recognition stage by an operator equipped with a sensor (e.g., a smartphone), who manually positions the sensor against an authentication region of a physical object. Registration also refers, in particular, to fine registration, that is, a precise alignment of images, based on detailed characteristics such as contours, textures, or points of interest, and which can be down to the pixel or even sub-pixel level. It should be noted that in the context of coarse registration, the implementation of the residual transformation is not always necessary.It should also be noted that the concept of registration, within the scope of the invention, covers the registration of two images relative to each other, but also the registration of each of two images relative to the same third image, which could be called the pivot image. Thus, within the scope of the invention, the images to be processed can all be registered relative to the same pivot image.
[0089] Following the first module 30, the input module INPUT includes a second module 31 which performs a superposition of the reference and candidate images after their processing by the first module 30. The superposition image from the second module 31 is then provided to the slicing module ECH. According to this example, the slicing occurs after the preliminary operations performed on the reference and candidate images, and just before the resulting imagelets (img) are provided to the neural network CNN. However, this order is not strictly necessary, and other sequences of the registration, residual transformation, and slicing steps can be considered. Thus, the slicing could occur after the recalibration but before the application of the residual transformation which would then be applied to the pairs resulting from the splitting.
[0090] Fig. 2 illustrates the result of the clipping of the superposition of two images which have undergone registration and residual transformation before their clipping, while Fig. 3 illustrates the result of the superposition of the imagelets from the same two images which have undergone registration then clipping and finally a residual transformation applied to the imagelets.
[0091] According to this first embodiment, the neural network is trained with two balanced training sets: a set of "recognized" image sets consisting of image sets resulting from the superposition of image sets containing at least one texture with a random component, this superposition exhibiting the Glass phenomenon, and a set of "unrecognized" image sets consisting of image sets resulting from the superposition of image sets containing at least one texture with a random component, this superposition not exhibiting the Glass phenomenon. The sets are said to be balanced in that they comprise the same number, or at least a very close number, of image sets. In this case, each training set comprises at least 1000 image sets and, preferably, more than 10,000 image sets.
[0092] Figure 4 shows an example of an imagelet imgl belonging to the set of recognized labeled images and exhibiting a Glass pattern. This imagelet imgl is the result of superimposing the two images located to its left on the same line, which form a pair PI of training images, labeled "recognized". Figure 4 also shows an example of an imagelet belonging to the set of "unrecognized" images and not exhibiting a Glass pattern. This latter imagelet is the result of superimposing the two images located to its left on the same line, which form a pair P2 of training images, labeled "unrecognized".
[0093] In a second embodiment more particularly illustrated in [Fig.5], the candidate and reference images are processed without implementing a superposition of the latter prior to their provision to the CNN neural network.
[0094] The CNN' neural network implemented in this second embodiment is a model analogous to the CNN described in relation to [Fig. 1]. It differs in that the image img' processed by the CNN' network will be twice the size of the image img processed by the CNN network. Thus, in this second embodiment, the image img' will have a size of m' x n' pixels, with, for example, one of the two m' or n' greater than 128 pixels and the other greater than 64 pixels. In this case, the image will have dimensions of 128 by 256 pixels.
[0095] The reference and candidate images will generally have a dimension greater than that of the imagelet img', so the ECH slicing module of the input block is configured to slice the candidate image and the reference image into PAV tiles, each of which has a size and shape corresponding exactly to half of the imagelet img'. As shown in [Fig. 5], the ECH slicing module slices the reference image IMG ref and the candidate image IMG cand into PAV tiles in a common coordinate system so as to provide at the output of the slicing the set of pairs of corresponding tiles; that is to say, all the pairs, each consisting of a tile of the reference image PAV ref and a tile of the candidate image PAV cand having exactly the same coordinates in the coordinate system as those of the tile PAV ref.
[0096] According to the illustrated example, the tiles of each pair will be concatenated, juxtaposed, to constitute each image Img' provided to the neural network CNN'. This operation will be performed, according to the illustrated example, by the slicing module but could also be performed by another module distinct from the slicing module. In the context of the invention and the present application, the term pair refers either to the result of the concatenation of two tiles, two images, or two imagelets, or to the two tiles, two images, or two imagelets in that they are intended to be processed jointly without having been assembled or concatenated.
[0097] In the second embodiment, the OUTPUT block will have the same mode of operation as that of the embodiment described in relation to [Fig.1].
[0098] According to a first variant of this second embodiment, the neural network is trained with two balanced training sets: a first set of images labeled "recognized" and a second set of images labeled "not recognized". Each image in the first set is formed by the concatenation of two image tiles comprising at least one texture with a random component, the two tiles being capable, after superposition, of exhibiting the Glass phenomenon. The PI pair in [Fig. 4] shows two tiles which, when concatenated, form one image from the "recognized" set.As previously indicated by "tiles capable of producing Glass patterns," this refers to tiles whose superposition would reliably produce a Glass pattern for a human operator with normal visual acuity, possibly after some registration operations, such as manual registration. Within the first training set, the tiles constituting the same training image can, upon simple superposition without any further operation, either produce a Glass pattern or not produce a Glass pattern for an observer with normal visual acuity.
[0099] Each image in the second set is formed by concatenating two image tiles, each containing at least one texture with a random component. The superposition of the two tiles does not exhibit the Glass phenomenon, even after fine registration followed by the application of a residual transformation. Pair P2 in [Fig. 4] shows two tiles which, when concatenated, form an image from the "unrecognized" set.
[0100] The CNN' neural network, thus trained, achieves a score of around 80% during the operational phases, with no significant improvement possible even by increasing the size and number of training sets. It should be noted that, in the first variant of the second embodiment of the invention, and according to the example described above, the input block comprises only the ECH slicing module, and no registration or residual transformation is performed on the candidate and reference images prior to their slicing. However, it is possible to perform registration and / or residual transformation on the reference and candidate images prior to their slicing.
[0101] According to a second variant of the second embodiment of the invention, the same CNN' neural network is implemented as before, but its learning, training, is carried out with differently constituted training sets.
[0102] Thus, according to the second variant of this second embodiment, the neural network is trained with two balanced training sets: a first set of images labeled "Glass-direct" and a second set of images labeled "non-Glass".
[0103] Each "Glass-direct" image in the first set consists of the concatenation of two image tiles comprising at least one texture with a random component, the two tiles exhibiting, after superposition, the Glass phenomenon for an observer with normal visual acuity without the need for any prior processing or operation, or even during the superposition. This will be referred to as simple superposition.
[0104] Each "non-Glass" image in the second set is made up of the concatenation of two image tiles comprising at least one texture with a random component, the superposition of the two tiles not being able to cause the Glass phenomenon even after a fine registration operation followed by the application of a residual transformation.
[0105] Once the learning has been carried out with the two sets "Glass direct" and "non-Glass", the CNN' neural network obtains in the exploitation phase a success rate of more than 90%.
[0106] Training according to the second variant makes it possible to obtain a CNN' neural network less sensitive to image acquisition and registration conditions, which makes it more robust and facilitates its implementation in the context of industrial and / or consumer applications for the authentication of mass-produced material subjects.
[0107] In the preceding example, two categories or classes of imagelets were identified: "Glass-direct" imagelets and "non-Glass" imagelets. It is possible to identify a third category of imagelets called "Glass-indirect." Each "Glass-indirect" imagelet consists of the concatenation of two image tiles, each containing at least one texture with a random component. The simple superposition of these two tiles does not exhibit the Glass phenomenon for an observer with normal visual acuity. In contrast, the superposition of the two tiles of the "Glass-indirect" imagelet exhibits a Glass pattern, either after coarse registration or after fine registration followed by a residual transformation. The imagelets of the specific categories "Glass-direct" and "Glass-indirect" both belong to the general category of imagelets "capable of generating the appearance of Glass patterns."
[0108] In the context of this application, it may be said that the tiles of a "Glass Direct" image are normalized or normalized with respect to each other. This normalization can be carried out by an operator or semi-automatically with automatic registration and operator control of the overlay, or completely automatically with automatic registration and the automatic application of a predefined residual transformation.
[0109] Surprisingly, the CNN' neural network trained according to the second variant of the second embodiment proves capable of identifying certain "glass-indirect" pairs or imagelets as "recognized," meaning that it is able to identify non-normalized imagelets or pairs of imagelets even though the transformation from one to the other is outside the domain of residual transformation. This characteristic is particularly advantageous because it allows, in practice, the implementation of a coarse registration, which is faster and less resource-intensive than a fine registration.Similarly, the CNN' neural network trained according to the example of the second variant is able to classify in the recognized category a pair of identical images which belong to the "Glass-indirect" category insofar as their simple superposition cannot reveal a Glass pattern but it is possible to reveal the Glass pattern by superimposing the two images one of which will have undergone a residual transformation.
[0110] When the CNN' neural network thus trained is put into operation, the input block preferably comprises, upstream of the ECH slicing module, A registration module 40 ensures the registration of the reference (IMG) and candidate (Cand) images relative to each other. The INPUT block then includes a transformation module 41 that induces a residual transformation relative to each other on the images supplied to it. For this purpose, the transformation module 41 applies, for example, the residual transformation to one of the two images and not the other. However, the transformation module 41 could operate differently, provided that the final result of the processing it performs corresponds to a residual transformation relative to each other of the two images supplied to it. As an example, the residual transformation is a rotation of a few degrees combined with a rectilinear translation of about ten pixels. According to the example illustrated [Fig. 5], the transformation module 41 is located between the registration module 40 and the slicing module ECH.However, it could be considered to place the transformation module 41 after the ECH slicing module. The transformation module 41 then works on the reference tiles Ref and candidate Cand before their concatenation, which is performed by a concatenation module not shown.
[0111] In the context of the invention and of the present application, the terms images, imagettes and tiles refer to objects, in the mathematical and computer science sense of the term, of the same nature, so that a treatment described in the context of the present application as being applied to one of these three types of object can, mutatis mutandis, be applied to the other two objects depending on the time when said treatment is implemented in the course of the process according to the invention.
[0112] According to a third embodiment of the invention, more particularly illustrated in [Fig. 7], at least two artificial neural networks, CNN1 and CNN2, are implemented. Each network has the same structure as the CNN and CNN' networks, with the sole difference that the LOSS module is common to both CNN1 and CNN2. Each CNN1 and CNN2 network processes an image from a pair, such that each image from the same pair is processed by a separate artificial neural network. The LOSS module calculates, for example, the Euclidean distance between the result of the processing performed by CNN1 and the result of the processing performed by CNN2. The results from CNN1 and CNN2 correspond to descriptors of each image introduced into each network, respectively.
[0113] The two artificial neural networks operate in parallel, and during operation, the first network, CNN1, processes the imagelets img ref from the reference image IMG ref, and the second network processes the imagelets img cand from the candidate image IMG cand. The two networks, CNN1 and CNN2, operating in parallel, can be called a Siamese network.
[0114] The training of the Siamese network thus constituted is carried out using the same training sets as those described in relation to the second and third variants of the second embodiment of the invention, with the difference that the tiles are not concatenated and form the image tiles of the training pairs that are provided to each of the CNN1 and CNN2 networks constituting the Siamese network. During training, the LOSS function must tend towards a zero value for the "recognized" or "Glass-direct" sets.
[0115] The operation of this third embodiment of the invention is carried out in a manner analogous to that described in relation to the second variant of the second embodiment. In this respect, it should be noted that the INPUT block has the same mode of operation except with regard to the concatenation of the tiles before their supply to the Siamese network, while the OUTPUT block has the same operation.
[0116] According to a variant of the third embodiment of the invention, more particularly illustrated in [Fig.8], three artificial neural networks CNN1, CNN2 and CNN3 are implemented in the learning phase, each of which has the same model as the CNN and CNN' networks, with the sole difference that the LOSS module is common to the three networks CNN1, CNN2 and CNN3.
[0117] The set of three CNN1, CNN2 and CNN3 networks operating in parallel can also be called a Siamese network the learning phase using a triplet of images, two of which represent elements of the class "recognized images" and one of the class "not recognized" always with the aim of minimizing the LOSS function, and which allows better discrimination performance in running such a type of CNN neural network.
[0118] In the context of the examples described above, the labels used to distinguish the training sets are "recognized", "not recognized", "Glass-direct", "Glass-indirect" and "Non-Glass" to facilitate the disclosure of the invention, but other labels could be used insofar as they allow the types of training sets to be distinguished from each other.
[0119] Preferably, the image sets used in the training sets include a large number of image sets with random textures from different types of material. This configuration of the training sets results in a more efficient CNN neural network, as it is capable of identifying Glass patterns on a large number of superimposed image sets with random textures from a wide variety of materials. Furthermore, the CNN thus trained will be able to recognize a Glass pattern resulting from the superposition of random textures that were not present in the training sets.
[0120] It should be noted that an artificial neural network exhibits, during operation, a behavior that reflects the learning stages it has undergone. Thus, an artificial neural network configured to be sensitive to Glass patterns in the context of images with random texture components will be recognizable among others, particularly by subjecting it to a set of tests containing such images and analyzing the responses provided by the artificial neural network under study.
[0121] Within the framework of the invention, the labeling of training images or image sets can be carried out in different ways. Initially, the labeling of the training sets is performed by an operator. This is referred to as supervised training. However, after an initial supervised training phase, it is possible to perform a second, unsupervised automatic training phase in which the image labeling is carried out by a CNN neural network that has undergone initial supervised training.
[0122] Similarly, the image sets belonging to the training sets can be derived from textured images with a random component of physical subjects, called natural image sets, or be image sets resulting from the superposition of synthetic images, called synthetic image sets. According to the invention, the training sets can comprise only natural image sets, only synthetic image sets, or a mixture of natural and synthetic image sets. Synthetic images are very well suited to simulating different lighting conditions (orientation, types, etc.), which can be very powerful for preparing a CNN for different shots of the candidate image and increasing its robustness.
[0123] By way of example, the applicant advantageously used randomly generated Perlin texture images to construct the training sets of a CNN' neural network according to the invention which, after training, proved effective for the authentication of paper-type material subjects.
[0124] Furthermore, training sets can be created by one or more operators through manual or supervised semi-automatic operations. However, the creation of training sets could be fully automated. Thus, training sets of imagelets or imagelet pairs of the "Glass-direct" type can be automatically generated from series of images with a continuous random component texture of various material subjects. The generation of each "Glass-direct" imagelet pair consists of extracting an imagelet from an image in said series, which forms the first image of the pair, and the second imagelet of the pair is formed by the first image after a residual transformation has been applied. For "Glass-indirect" pairs, the procedure can be the same as for "Glass-direct" pairs, with the difference that instead of applying a residual transformation to the first image to create the second image, a rotation between 90° and 270° is applied to the first image to form the second image.
[0125] Similarly, training sets of imagelets or imagelet pairs of "non-Glass" type can be automatically generated from series of images with a continuous random component texture of various material subjects. The generation of each pair of "non-Glass-direct" imagelets consists of extracting an imagelet from a textured image of two distinct subjects.
[0126] Furthermore, within the scope of this application and the context of the invention, the terms blocks and modules may refer to purely software elements, purely hardware elements, or elements combining software and hardware implementations. This also applies to the implementation of the entire invention, which may be purely software-based, purely hardware-based, or a combination of dedicated software and hardware components.
[0127] It should also be indicated that within the framework of the invention the input block can implement various types of computational processes to carry out the different processing it performs, that it can in particular implement specific artificial neural networks configured to perform the operations of cropping, cutting and applying transformations to the images to be processed.
[0128] Furthermore, in the examples described above, the images to be processed are segmented before being provided to each neural network configured to be sensitive to Glass patterns. However, such a mode of operation is not strictly necessary for the realization of the invention, since it is possible to implement artificial neural networks sized for processing images from acquisitions without it being necessary to segment these images beforehand.
Claims
Demands
1. A method for authenticating a material subject consisting of confronting, on the one hand, at least one so-called reference image of at least one authentication region of an authentic subject, the reference image comprising at least one texture with a random component, and on the other hand, a so-called candidate image of at least one authentication region of a candidate subject, the candidate image comprising at least one texture with a random component, characterized in that it implements at least one artificial neural network configured to be sensitive to Glass patterns in the context of images with a texture with a random component.
2. Authentication method according to claim 1, characterized in that the artificial neural network is provided with an image resulting from the superposition of the candidate image and the reference image, the artificial neural network having been trained using images with random component texture exhibiting a Glass pattern.
3. Authentication method according to claim 2, characterized in that prior to their superposition the candidate and reference images are re-registered with respect to each other.
4. Authentication method according to claim 3, characterized in that, after registration and before their superposition, the candidate and reference images are subjected to a residual transformation with respect to each other.
5. Authentication method according to claim 1, characterized in that the candidate image and the reference image are provided to the artificial neural network without superimposition of these, the artificial neural network having been trained by means of pairs of non-superimposed random component textured images whose superimposition reveals a Glass pattern, called training pairs.
6. Authentication method according to claim 5, characterized in that the images of each training pair are calibrated against each other prior to their submission to the artificial neural network in the training phase.
7. Authentication method according to claim 6, characterized in that subsequent to the calibration and prior to supply In the artificial neural network, the images of each training pair undergo a residual transformation relative to each other.
8. Authentication method according to any one of claims 5 to 7, characterized in that prior to their provision to the artificial neural network the candidate and reference images are calibrated against each other.
9. Authentication method according to claim 8, characterized in that, after registration and before being supplied to the artificial neural network, the candidate and reference images are subjected to a residual transformation with respect to each other.
10. Authentication method according to any one of claims 5 to 9, characterized in that just before being supplied to the artificial neural networks, the two images to be processed, either from the training pair, or the reference image and the candidate image, are assembled by being juxtaposed without overlap to form only one and the same image which will be supplied to the artificial neural network.
11. Authentication method according to claim 1 characterized in that it implements two twin artificial neural networks operating in parallel, the reference image being provided to a first network and the candidate image being provided to the other network, the two networks having been trained by means of pairs of non-superimposed random component textured images whose superposition reveals a Glass pattern, called training pairs, the two images of the same pair each being provided to a separate artificial neural network.
12. Authentication method according to claim 11, characterized in that, prior to their provision to the artificial neural networks, the images of each training pair are recalibrated against each other.
13. Authentication method according to claim 11 or 12, characterized in that, after registration and before being supplied to artificial neural networks, the images of each training pair are subjected to a residual transformation with respect to each other.
14. Authentication method according to any one of claims 11 to 13, characterized in that the reference and candidate images are recalibrated against each other before being supplied to artificial neural networks.
15. Authentication method according to claim 14, characterized in that, after registration and before being supplied to neural networks, the candidate and reference images are subjected to a residual transformation with respect to each other.
16. Authentication method according to any one of the preceding claims, characterized in that each artificial neural network is adapted to process images or imagelets of given dimensions and when the images to be processed are larger, the images to be processed are cut into sub-images or imagelets of suitable size which are successively submitted directly to each artificial neural network.
17. Authentication method according to any one of the preceding claims characterized in that each candidate image is taken from a video stream.
18. Authentication method according to any one of the preceding claims, characterized in that at least one artificial neural network is a convolutional neural network.
19. Product computer program comprising code instructions for performing a method according to any one of the preceding claims of authenticating a material subject from a reference image and a candidate image, when the program is run on a computer.
20. A computer-readable storage means on which a computer program includes code instructions for performing a method according to any one of claims 1 to 18 of authenticating a material subject from a reference image and a candidate image.
21. Computer device comprising at least display means, image acquisition means, user input means, information storage means communicating with computing and control means configured to implement the method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Method for authenticating and / or checking the integrity of a subject
EP3380987A1
Method of augmented authentification of a material subject
US10990845B2
Non-counterfeitable document system
US4423415A
procedure D'AUTHENTIFICATION PAR MOTIF DE GLASS
FR3044451A3