Method for authenticating material subjects using a glass pattern by means of a siamese neural network
Siamese neural networks with image registration and residual transformations enable efficient and accurate authentication of material subjects with limited resources, addressing the limitations of existing methods.
Patent Information
- Application Number
- PCT/EP2025/072835
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-08
- Filing Date
- 2025-08-07
- Publication Date
- 2026-02-12
AI Technical Summary
Existing methods for authenticating material subjects require significant computing power and long computation times, which is impractical for real-time applications with limited resources and high energy consumption.
A method using Siamese artificial neural networks, specifically convolutional neural networks, to authenticate material subjects by comparing reference and candidate images with random textures, employing training phases that include image registration and residual transformations to enhance authentication accuracy.
Achieves high authentication rates of around 95% using limited computing resources and reduced energy consumption, allowing for flexible image acquisition conditions and robust authentication despite minor modifications to the subject.
Smart Images

Figure EP2025072835_12022026_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR AUTHENTICATING MATERIAL SUBJECTS BY GLASS PATTERN USING A SIAMESE ARTIFICIAL NEURAL NETWORK
[0002]
[0001] The present invention relates to the technical field of unitary authentication of material subjects with electronic means.
[0003]
[0002] In the above-mentioned field, it is known, notably from US patent 4,423,415, to identify physical objects by extracting a signature from a so-called authentication region comprising an essentially random, three-dimensional intrinsic microstructure. Initially, a first extraction, of a so-called reference signature of an authentic object, is generally performed using electronic computing tools such as a computer, smartphone, or tablet, which implement complex pattern recognition algorithms. The reference signature is then recorded. Subsequently, a second extraction, of a so-called candidate signature of a candidate object, is performed using electronic computing tools similar to those used in the first step.The reference and candidate signatures are then compared with similarity or correlation comparison algorithms to determine their level of proximity and to deduce, if applicable, that the candidate subject is indeed the authentic subject.
[0004]
[0003] US patent 10990 845 proposed a method for unit authentication of physical objects by calculating similarity vectors between, in particular, reference and candidate images of authentication regions of authentic and candidate physical objects. The methods and algorithms implemented in this patent require significant computing power for real-time implementation or long computation times with less power if high execution speed is not required. However, this constraint of significant computing power or long computation times proves to be a disadvantage when it is necessary to obtain computation times as short as possible while having limited computing power combined with low energy consumption.
[0005]
[0004] It therefore arose the need for alternative solutions which provide a solution to this problem and allow for reliable unit authentication of material subjects while using limited computing resources and contained energy consumption and which offer greater flexibility with regard to the conditions of acquisition of reference and candidate images than prior art processes.
[0006]
[0005] To achieve this objective, the invention relates to a method for authenticating a material subject consisting of comparing, on the one hand, at least one so-called reference image of at least one authentication region of an authentic subject, the reference image comprising at least one texture with a random component, and on the other hand, a so-called candidate image of at least one authentication region of a candidate subject, the candidate image comprising at least one texture with a random component. According to the invention, comprises:
[0007] - in a training phase:
[0008] - the implementation of a Siamese set comprising at least two twin artificial neural networks that are adapted to operate in parallel and whose outputs are connected during the learning phase to the same loss stage implementing a loss function,
[0009] - a training step in which the Siamese set is trained with at least one training set comprising training pairs, at least some of which comprise two images whose superposition reveals or is likely to reveal a pattern analogous to a Glass pattern, the two images of the same training pair each being provided to a separate artificial neural network belonging to the Siamese set
[0010] - after the training phase, in an operational phase:
[0011] - the implementation, without the loss stage, of a twin network of artificial neurons belonging to the Siamese set trained to be sensitive to Glass patterns in the context of images with random texture components,
[0012] - a step of providing the twin network with the reference image so that the twin network can calculate a descriptor of the reference image,
[0013] - a step of providing the twin network with the candidate image so that the twin network can calculate a descriptor of the candidate image,
[0014] - a step of calculating a degree of similarity between the descriptor of the reference image and the descriptor of the candidate image.
[0015]
[0006] The degree of similarity can then be compared to a threshold above which the authentication result is positive and below which there is no authentication. Such a descriptor can also be used in the context of the method for relating a candidate image to a reference image described in patent application EP4396789.
[0016]
[0007] According to the invention, during the operational phase, both Siamese networks can be implemented without the loss stage. However, in a preferred embodiment, only one twin network is implemented during this operational phase; implementing a single twin network of artificial neurons during the operational phase allows for very good performance with limited resource usage.
[0017]
[0008] In a preferred embodiment, the invention implements at least one convolutional type artificial neural network.
[0018]
[0009] For the purposes of this invention, a convolutional artificial neural network, also called a convolutional artificial neural network, is a type of neural network comprising at least one and, preferably, several convolutional layers. Hereafter, the terms "convolutional artificial neural network," "convolutional neural network," and "convolutional neural network" shall be used interchangeably as synonyms. Similarly, within the scope of this invention, the terms "neural network" and "artificial neural network" are synonymous, and the term "neuron(s)" shall be understood to mean "artificial neuron(s)" unless otherwise specified.
[0019]
[0010] The implementation of convolutional artificial neural networks makes it possible to obtain satisfactory computation times with reasonable computing resources, such as those available in portable devices like smartphones or tablets, while exhibiting lower energy consumption than that required by prior art computational methods. Furthermore, the implementation of artificial neural networks allows for a certain flexibility in the acquisition conditions of reference and candidate images.
[0020]
[0011] For the purposes of the invention, a convolutional artificial neural network configured to be sensitive to Glass patterns in the context of images with random texture components is understood to be an artificial neural network or a Siamese neural network that has been trained at least with images whose superposition reveals or is likely to reveal a pattern analogous to a Glass pattern, as described in EP 3,380,987 patent filed by the applicant, to which reference should be made. Within the scope of the invention, the artificial neural network(s) implemented are configured to respond to images that, in combination, reveal or are likely to reveal said Glass pattern, or to an image resulting from the superposition of two images and revealing said Glass pattern.
[0021]
[0012] Thus, the present invention takes advantage of the applicant's demonstration that patterns similar to those obtained by Léon GLASS in articles in the journal NATURE vol.223 of August 9, 1969 pages 578 to 580 and Nature vol. 246 of December 7, 1973 pages 360 to 362, can appear by superimposing two images comprising respectively natural textures or random components resulting from the acquisition or even the photography at an appropriate magnification or enlargement of the same intrinsic random multi-scale three-dimensional material structure of the same subject.The applicant has demonstrated that these Glass-type patterns appear only when continuous random textures originating from the same material structure and essentially residual geometric transformations of each other are superimposed. In practice, they do not appear when the continuous random textures are not sufficiently correlated or do not result from the acquisition of the same material structure corresponding to a subject's recognition region. Therefore, within the scope of the invention, Glass patterns enable unitary authentication, also known as unitary recognition. Conversely, if a Glass pattern is not observed, it is not possible to definitively conclude that the data is not authentic.Furthermore, the applicant highlighted the fact that, in the case of physical subject authentication, if said physical subject exhibits sufficient stability over time, images taken at different times, even if separated by several days, months, or years, can, through their superposition, generate such Glass patterns. Moreover, according to the invention, the authentic subject can undergo modifications after the authentication image is recorded while remaining authenticable, provided that a portion of the recognition region has not been significantly affected by these modifications, whether intentional or not.Furthermore, the applicant has demonstrated that the observation of Glass patterns is possible by superimposing two images of the same recognition region possessing a texture with a continuous random component, without the addition, extraction, or generation of discrete elements or discrete patterns as advocated in the prior art before EP 3,380,987 in the applicant's name. Such a Glass pattern appears only in the case of an authentic subject and only if there is a slight non-zero geometric transformation, referred to in the sense of the invention as a residual geometric transformation, between the candidate image and the authentication image, or between the verification and authentication images acquired under the given conditions.This property of Glass patterns provides the method according to the invention with great robustness, since it is not necessary for the acquisition conditions of the verification image to be strictly identical to the acquisition conditions of the authentication image. Thus, the resolutions of the authentication and verification images may differ. In the context of this application, an authentication image is also referred to as a reference image, while a verification image is also referred to as a candidate image.
[0022]
[0013] Within the framework of the invention, an authentication region, also called a recognition region, is a region that exhibits an intrinsic and random microstructure in that it results from the very nature of the authentication region of the material subject. In a preferred embodiment of the invention, each material subject used belongs to the families of subjects comprising at least one authentication region having an essentially random intrinsic structure that is not easily reproducible, that is to say, whose reproduction is difficult or even impossible in that it results in particular from a process that is not predictable at the observational scale.Such an authentication region with an essentially random, non-easily reproducible intrinsic continuous medium structure corresponds to Physical Unclonable Functions (PUFs), as defined in particular by Jorge Guajardo in the English-language Encyclopedia of Cryptography and Security, 01 / 2011 edition, pages 929 to 934. Preferably, the authentication region of a material subject conforming to the invention corresponds to an intrinsic nonclonable physical function, designated as "Intrinsic PUFs" in the aforementioned article.The applicant takes advantage of the fact that the random nature of the authentication region's microstructure is inherent or intrinsic to the very nature of the subject, resulting from its process of formation, development, or growth. Therefore, it is unnecessary to add a specific structure to the authentication region, such as an impression or engraving. Regarding the recognition region, according to the invention, its image possesses a texture that an observer with average visual acuity can perceive either with the naked eye or via optically and / or digitally enhanced zoom. Thus, small structures, or those perceived as small at the observation magnification, are seen as images containing a texture.
[0023]
[0014] In the context of the invention, the term texture refers to what is visible or observable in an image, while the term structure or microstructure refers to the material subject itself. The texture of interest in the recognition region is described as a texture with a random component in that it includes at least some randomness or irregularity. Thus, a texture or microtexture of the recognition region, called a texture-with-a-random component image, corresponds to an image of the structure or microstructure of the recognition region. Imaging such a recognition region under similar viewing conditions from nearby viewpoints provides images, each containing a texture with a random component that is a noisy reflection of its material structure.Such a texture with a random component inherits its unpredictability and independence from a texture with a random component from a completely different recognition region, from the element of randomness in the formation of their material structures.
[0024]
[0015] For the purposes of this invention, the term "image," whether candidate, reference, or training, means any type of image in the general sense of the term, and not solely an image comparable to a photograph. In other words, an image is not limited to an optical image resulting from the stimulation of the recognition region by visible light, but can instead be obtained by any type of physical stimulation, including but not limited to: ultrasound, far-infrared, terahertz, X-rays or gamma rays, X-ray or laser tomography, X-ray radiography, and magnetic resonance imaging. Thus, for the purposes of this invention, an image is, for example, the recording of the result of stimulation by any means whatsoever of a natural scene or a material subject. This recording can then be described as a natural image.This recording may be one-dimensional, corresponding, for example, to the recording of the variation over time or along a line of a single signal, or to the recording of the values from a line of sensors. This recording may also be two-dimensional, as is the case with a photograph, which can be recorded in halftones, grayscale, or color. For the purposes of this invention, an "image" can therefore be a 1D signal, a 2D or 3D image, in grayscale or color, or an nD (n-dimensional) signal, for example, hyperspectral or RGB-D. For the purposes of this invention, an "image" can be the result of a single acquisition or extracted from a stream. Thus, within the scope of this invention, an "image" can be extracted from a video stream. In one embodiment, the image can be saved digitally.Furthermore, an image used in the process according to the invention can be a pre-segmented image, particularly when it comprises several physical subjects in the same scene or a repetition of the same motif. Of course, an image is not necessarily natural and can be synthetic, meaning it can be generated by a computer process with or without the assistance of a human operator. Within the scope of the invention, a natural image and a synthetic image have in common that they are in the same digital or analog recording format for processing within the same process. Optical and / or digital image enhancement pre-processing can also be applied to it, for example, to improve the signal-to-noise ratio.
[0025]
[0016] In the context of the invention, training, also called learning, is carried out by means of, on the one hand, training pairs comprising two images whose superposition is known to generate a Glass pattern, without any special precautions regarding their presentation to the artificial neural network to be trained, and, on the other hand, training pairs comprising two images whose superposition is known not to generate Glass patterns as defined in the invention. By "no special precautions," it should be understood that the images in the same pair have not been subjected to any registration, whether during the training or processing phase.
[0026]
[0017] By proceeding in this manner, a Siamese pair configured to be sensitive to the Glass pattern was obtained, achieving an identification rate of approximately 80% during operation, without it being possible to substantially improve this authentication rate by increasing the size of the training datasets. By "images whose superposition is known to be likely to generate Glass patterns," it should be understood that these are images from the superposition of which a human operator will certainly observe a Glass pattern, possibly after registration operations have been performed, for example, manually.
[0018] Such an identification rate of 80% may be satisfactory in certain applications, in particular, when it is possible to have a second check performed by an operator implementing the invention as described in patent EP 3 380987 in the name of the applicant.However, this rate is not sufficient when it is necessary to process a large number of material subjects, or their images, in a limited and reduced time, limiting the possibility of human intervention.
[0027]
[0019] The need therefore arose to improve the authentication or recognition rate. To this end, a variant of the training phase of the method according to the invention proposes to register the images of each training pair with respect to each other prior to their submission to the artificial neural networks during the training phase. It was found that, quite surprisingly, implementing registration and then applying a residual transformation prior to presenting the images of the training pairs yielded better results than when no precautions were taken. By proceeding in this way, it was possible to achieve authentication rates of around 95% or even higher during the operational phase.
[0028]
[0020] According to a feature of this variant, after registration and before supplying to the artificial neural network, the images of each training pair are subjected to a residual transformation with respect to each other.
[0029]
[0021] For the purposes of this invention, a residual transformation or residual geometric transformation is a slight, non-zero geometric transformation to be applied locally to the chosen image or images, whether rigid or not, linear or non-linear, at at least one fixed or quasi-fixed point. Among the applicable geometric transformations, it is thus possible to implement the transformations described by Leon Glass in his 1973 and 2002 articles cited above. A quasi-fixed point is understood to be a point that, after the residual geometric transformation, undergoes a displacement of small magnitude compared to the maximum displacement caused by the residual geometric transformation.In the theoretical case of a perfect superposition of strictly identical elements / images, there is no appearance of a Glass pattern even in the presence of an authentic subject, hence the need for the presence of this residual geometric transformation and the general interest in implementing a movement or relative displacement or even a deformation induced by a difference in shooting angle or viewpoint between the acquisitions of the authentication image and the verification image.
[0030]
[0022] According to a feature of the invention, aimed at increasing the performance of the authentication process, the training set comprises training pairs comprising two distinct images of the same texture with a random component with different shooting conditions and whose superposition reveals or is likely to reveal a pattern analogous to a Glass pattern.
[0031]
[0023] Among the parameters defining the conditions for capturing an image, the following may be mentioned: the image acquisition sensor, the lens associated with said sensor and the adjustment of this lens, the shooting angle, and the lighting conditions during shooting, without this list being limiting or exhaustive. In the context of the invention, the shooting conditions differ insofar as at least one of these parameters is not identical for the two images, while of course not preventing the appearance of the Glass pattern in the event of superimposition of the two images after registration and / or residual transformation, if applicable.
[0032]
[0024] According to one feature of the invention, during the operating phase, the reference image descriptor is recorded and the candidate image is provided to the twin network after this recording. Thus, according to this feature, and in the context of, for example, the authentication of manufactured objects, the reference image descriptors of these objects are recorded at their place of production before they leave that place of production. When authentication of one of these objects is desired after they have left their place of production, a candidate image is acquired, and this candidate image is then processed according to the method of the invention to obtain a descriptor that will be compared to the previously recorded reference image descriptors.In the context of the invention, a descriptor can be of vector, matrix or even tensor type, and obtained for example after the execution of at least one convolution layer, the execution of other layers (e.g. FDC) being optional.
[0033]
[0025] According to another feature of the invention, during the exploitation phase, the candidate image is registered with respect to the reference image before being provided to the artificial neural twin network.
[0026] According to a variant of this feature, after registration and before being provided to the artificial neural twin network, the candidate image undergoes a residual transformation with respect to the reference image.
[0034]
[0027] According to another feature of the invention, during the exploitation phase, the candidate image and the reference image are registered with respect to the same pivot image before being supplied to the twin artificial neural network. The implementation of such a pivot image makes it possible to perform registration of the candidate image relevant to the process according to the invention without it being necessary to have the reference image.
[0035]
[0028] According to a variant of this feature, after registration and before being supplied to the twin network of artificial neurons, the candidate image and / or the reference image is subjected to a residual transformation with respect to the pivot image.
[0036]
[0029] According to an embodiment of the method according to the invention, each artificial neural network is adapted to process images or imagelets of given dimensions and when the images to be processed are larger, the images to be processed are cut into sub-images or imagelets of suitable size which are successively submitted directly to each artificial neural network.
[0037]
[0030] In the context of this application and when referring to the processing performed by the invention, the term "image" is preferably used for the result of the acquisition, while the term "imagelet" is used for the object that is processed by each neural network. However, it should be noted that images and imagelets are of the same nature, differing only in size or dimensions, so that these terms may be used interchangeably depending on the context, without hindering the understanding of the invention by a person skilled in the art. Indeed, the terms "image(s)" and "imagelet(s)" are used to facilitate understanding of the description of the invention and should therefore not be interpreted as limiting.Similarly, in some cases, the sub-images may also be referred to as the block, this term also used to facilitate the description of the invention should not be interpreted in a restrictive manner.
[0038]
[0031] According to one feature of this embodiment of the method according to the invention, the candidate and reference images are registered prior to being clipped.
[0032] According to a preferred variant of this feature, the registration is performed before clipping, while the residual transformation is applied after clipping.
[0039]
[0033] The invention also relates to a computer program product comprising code instructions for executing a method according to the invention for the one-time authentication of a physical object from a reference image and a candidate image, when the program is executed on a computer. Such a computer program product can then be embedded or stored in a computer or similar device. For the purposes of this invention, the term "computer" should be understood in a broad sense, including, in particular, a personal computer, a smartphone, a tablet, a virtual or physical server, or any computing and processing unit capable of implementing a computer program and adapted for carrying out the invention.
[0040]
[0034] The invention also relates to a means of storage readable on a computer equipment on which a computer program includes code instructions for the execution of a method according to the invention of authenticating a material subject from a reference image and a candidate image.
[0041]
[0035] The invention also relates to a computer device comprising at least display means, image acquisition means, user input means, information storage means communicating with computing and control means configured to implement the method according to the invention of authenticating a material subject from a referenced image of a candidate image.
[0042]
[0036] The different modes of implementation, forms of embodiment, characteristics and variants of the invention can be implemented with each other in different combinations insofar as they are not mutually exclusive or incompatible with each other.
[0043]
[0037] Various other features and variants of the invention will become apparent from the description below, made in relation to the figures in which:
[0044] - Fig. 1 is a schematic representation of a Siamese system according to the invention implementing two twin artificial neural networks,
[0045] - Fig. 2 shows pairs of image sets of training sets used for training convolutional neural networks implemented within the framework of the invention, the right part of the figure showing the result of superimposing the image sets located on the same line on the left,
[0046] - Fig. 3 is a schematic representation of the operation of an image slicing module implemented within the framework of the invention,
[0047] - Fig. 4 illustrates the superposition of imagelets resulting from the processing of a reference image and a candidate image, in which the clipping is carried out after registration of the two images and application of a residual transformation,
[0048] - Fig. 5 illustrates the superposition of imagelets resulting from another processing of the reference image and a candidate image, in which the images are registered and then cropped, the residual transformation being applied to the imagelets obtained before their superposition,
[0049] - Fig. 6 is a schematic illustration of a preferred form of the execution phase of the process according to the invention.
[0050]
[0038] In the figures, the elements common to the different embodiments bear the same reference numerals. Furthermore, the different embodiments, presented in relation to the figures, correspond to non-limiting examples of possible executions and implementations of the invention.
[0051]
[0039] As previously stated, the invention implements a twin system comprising at least two convolutional neural networks to ensure the authentication of physical subjects from images of an authentication region of these subjects, the images comprising at least one texture with a random component. In a preferred embodiment, the invention makes it possible to ensure unitary authentication, that is, a reference image, also called an authentication image, corresponds to one and only one physical subject.
[0052]
[0040] Thus, unit authentication means unit recognition of a region of a material subject. This recognition can have a higher or lower probative value depending on the criticality of the use case considered and the measures implemented to increase this probative value, such as, but not limited to: the involvement or not of a trusted third party, the selection of highly sophisticated acquisition sensors in terms of resolution and illumination conditions for acquiring reference images and the use of the same sensors for acquiring candidate images, and complete control of the IT environment implemented, without this list being exhaustive or limiting.
[0053]
[0041] According to the invention, the authentic subject can undergo modifications after the recording of the reference image, also called the authentication image in the context of patent EP 3 380 987, while remaining authenticable to the extent that a part of the authentication region has not been profoundly affected by these voluntary or involuntary modifications.
[0054]
[0042] In order to be implemented on portable devices and / or to avoid requiring significant computing resources, which are also energy-intensive, the invention proposes implementing artificial neural networks that could be described as frugal in that they comprise a limited or even reduced number of layers. In a preferred embodiment, the invention implements convolutional neural networks comprising a limited number of convolutional neuron layers and fully or densely connected neuron layers.
[0055]
[0043] Among the examples of convolutional neural networks, also known by the abbreviation "CNN" (Convolutional Neural Network), that could be implemented by the invention, we can cite LetNet, LetNeT-5, or AlexNet neural networks. Reference can also be made to the French and English Wikipedia pages entitled "convolutional neural network" and
[0056] "convolutional neural network" to have other examples of convolutional neural networks and explanations of their structure.
[0057]
[0044] According to a preferred embodiment of the invention, at least two identical convolutional neural networks are used, each comprising a stack of processing layers, namely: a. convolutional layers (CONV) which process the data from a receiving field, namely an image, b. pooling layers (POOL) which allow the information to be compressed by reducing the size of the intermediate image, c. correction layers, often incorrectly called ReLU by reference to the rectified linear activation function, d. completely or densely connected layers (FDC), e. a LOSS layer, which can also be called the loss layer common to all the neural networks constituting the Siamese system.
[0058]
[0045] It should be noted that for some authors the LOSS layer is considered not to be part of the neural network; therefore, within the scope of the invention, a neural network does not necessarily have such a layer. Furthermore, the so-called LOSS layer is present only during the learning phase and not during the execution phase. Within the scope of the invention, the terms "loss layer," "loss stage,"
[0059] "Loss floor" are synonymous. Similarly, the terms "floor" and "layer" are equivalent.
[0060]
[0046] Thus, and as can be seen from Figure 1, the Siamese system comprises two twin convolutional artificial neural networks, designated CNN1 and CNN2 respectively. Since the twin networks CNN1 and CNN2 are identical, they will also be generically referred to by the acronym CNN.
[0061]
[0047] The Siamese system is adapted to process pairs of images; the two networks CNN1 and CNN2 each process one image from a pair, such that each image in the same pair is processed by a distinct artificial neural network. The LOSS module or stage calculates, for example, the Euclidean distance between the result of the processing performed by the CNN1 network and the result of the processing performed by the CNN2 network. The results of the CNN1 and CNN2 networks correspond to descriptors of each image introduced into each network, respectively.
[0062]
[0048] Each CNN twin network is configured to process a rectangular image img of mxn pixels, preferably with m and n greater than or equal to 64, for example, 128x128 pixels, it being understood that m and n are not necessarily equal. Each CNN neural network comprises a convolution processing block 1 followed by a completely or densely connected artificial neural processing block 2.
[0063]
[0049] According to the illustrated example, each CNN network further includes, between the convolution block 1 and the processing block 2 with completely or densely connected artificial neurons, a FLAT layer for processing the output of the convolution block 1 before supplying it to the block 2.
[0064]
[0050] According to the illustrated example, the convolution block 1 comprises successively and in this order:
[0065] - a first convolution layer 11, - a first correction layer 12,
[0066] - a first layer of pooling 13,
[0067] - a second convolution layer 14,
[0068] - a second correction layer 15,
[0069] - a second layer of pooling 16,
[0070] - a third convolution layer 17
[0071] - a correction layer 18 which happens to be the last layer of convolution block 1.
[0072]
[0051] The first convolution layer 11 is, in this case, a 2D convolution layer parameterized to process the entire image using 3x3 pixel tiles (kernels or filters) with a step of 1 pixel. The first correction layer 12 implements a Rectified Linear Unit (ReLU) activation function to process the result of the first convolution layer 11. The first pooling layer 13, as illustrated in the example, performs maximum pooling for 2D spatial data with a 2x2 window and a step of 2, on the result of the processing performed by the first correction layer 12.
[0073]
[0052] The second convolution layer 14 is, in this case, a 2D convolution layer parameterized to process the output of the first pooling layer 13 by 5x5 pixel tiles with a step of 2 pixels. The second correction layer 15 implements a Rectified Linear Unit (ReLU) activation function to process the result of the second convolution layer 14. The second pooling layer 16 ensures, according to the illustrated example, a maximum pooling operation for 2D spatial data with a 2x2 window and a step of 2 on the result from the second correction layer 15.
[0074]
[0053] The third convolution layer 17 is, in this case, a 2D convolution layer parameterized to process the result of the second pooling layer 16 by 5x5 pixel tiles with a step of 2 pixels. The third correction layer 18 here implements a ReLU (Rectified Linear Unit) type activation function to ensure the processing of the result of the third convolution layer 17. According to the illustrated example, the third correction layer 18 is the last layer of the convolution block 1.
[0054] An example of code for defining the convolution block 1 as described above in PyTorch is as follows: self.cnnl = nn. Sequential nn.Conv2d(l, 128, kernel_size=3, stride=l), nn.ReLU(inplace=True) , nn.MaxPool2d(2, stride=2), nn.Conv2d(128, 256, kernel_size=5, stride=2), nn . ReLU(inplace=True), nn .MaxPool2d (2, stride=2), nn . Conv2d(256, 512, kernel_size=5, stride=2), nn.ReLU(inplace=True) ,
[0075]
[0055] Each CNN neural network includes at the output of the convolution block 1 a FLAT processing layer, having the reference 19, which ensures the flattening in the form of a single-row vector or matrix of the result of the processing from the third correction layer 18.
[0076]
[0056] Downstream of the processing layer 19, each CNN neural network comprises the fully or densely connected neural network processing block 2. According to the illustrated example, block 2 comprises successively and in this order:
[0077] - a first layer of 20 densely or completely connected neurons
[0078] - a first layer of correction 21,
[0079] - a second layer 22 of densely or completely connected neurons,
[0080] - a second correction layer 23,
[0081] - a third layer 24 of densely or completely connected neurons which is, according to the illustrated example, the last layer of block 2.
[0082]
[0057] According to the illustrated example, the first layer 20 is a linear layer of artificial neurons receiving 512 inputs and delivering 1024 outputs. By linear layer, it is understood that a line of neurons, that is to say, a layer with a single thickness of neurons.
[0083]
[0058] The first correction layer 21, of block 2, here implements an activation function of type ReLU for "Rectified Linear Unit" applied to each of the 1024 outputs of layer 20.
[0059] The second layer 22 is a linear layer of artificial neurons receiving 1024 inputs and delivering 256 outputs.
[0084]
[0060] The second correction layer 23, of block 2, here implements an activation function of type ReLU for in English "Rectified Linear Unit" applied to each of the 256 outputs of layer 22.
[0085]
[0061] Finally, the third and last layer 24 is a linear layer of artificial neurons receiving 256 inputs and delivering 1 output.
[0086]
[0062] An example of code for the definition of block 2, for a fully or densely connected neural network, as described above in PyTorch language is as follows: self.fcl = nn.Sequential nn.Linear(512, 1024), nn.ReLU(inplace=True), nn.Linear(1024, 256), nn.ReLU(inplace=True), nn.Linear(256, 1),
[0087] )
[0088]
[0063] According to the illustrated example, the third and last layer 24 forms an output layer which, in this case, ensures a normalization of the value delivered by the third and last layer 23 of block 2, in the form of a real number, between 0 and 1.
[0089]
[0064] It should be noted that the different values of the functions mentioned in the Pytorch code examples are designated by the generic term hyperparameters; these are data which are not automatically updated during the learning phases, also called training phases, as opposed to parameters which are, such as, for example, the weight of the connections between artificial neurons and the biases of these.
[0090]
[0065] According to the illustrated example, an input block (INPUT) is implemented upstream of the CNN1 and CNN2 networks. This block subjects each image or pair of images to be processed to various processing operations before they are provided to the twin networks. For example, when the images to be processed (IMG) are larger than the image size (img) that can be processed by each CNN1 and CNN2 neural network, the input block (INPUT) will perform a segmentation or decomposition of each image to be processed (IMG) into sub-images or imagelets (img), which will be directly provided to each CNN neural network. In the context of this application, the segmentation performed by the segmentation module corresponds to a segmentation of each image into a multiplicity of smaller imagelets or into a multiplicity of smaller blocks suitable for processing by the CNN1 and CNN2 twin networks.
[0091]
[0066] The training phase, also called the learning phase, is carried out on the Siamese system, which comprises the twin networks CNN1 and CNN2 operating in parallel. The INPUT module is not generally required during the training phase, but its implementation is not excluded. This training phase includes a training step during which the Siamese system learns according to a number of operations well known to those skilled in the art, which are as follows:
[0092] 1 / Initialization of weights: The weights of the connections between neurons are initialized with small random values.
[0093] 2 / Presentation of training images: The images from each training pair are introduced, one at a time and in parallel, into one of the twin networks, starting with the input layer and propagating towards the output layer of the network. 3 / Output calculation: Each twin network calculates the output for the given input, applying nonlinear activation functions to each neuron.
[0094] 4 / Error calculation: The error between the actual output and the predicted output is calculated, usually using the loss function.
[0095] 5 / Backpropagation of the error: The error is propagated back through each twin network, across each layer, by adjusting the weights of the connections between neurons. 6 / Updating the weights: The weights are adjusted according to the error gradient, using an optimization method such as gradient descent.
[0096] 7 / Repetition: Steps 2 to 6 are repeated for a large number of epochs (iterations), until the Siamese system and therefore the twin networks that compose it reach an acceptable accuracy on the training images.
[0097] 8 / Evaluation / Test: The Siamese system is trained and evaluated on a set of test images to estimate its performance.
[0098]
[0067] An "epoch" in the context of machine learning corresponds to a complete pass through the twin network of all pairs of training (respectively, test) imagelets. In other words, an epoch is a complete iteration in which the model sees all pairs of training (respectively, test) imagelets once, and the weights and biases are updated accordingly. For example, if the set of training (respectively, test) imagelet pairs contains 1000 pairs of imagelets, an epoch corresponds to the presentation of these 1000 pairs of imagelets to the twin network, and the joint updating of the weights and biases of the twin networks after each batch of image pair(s). It is important to note that the term "epoch" is often used interchangeably with the term "iteration," but they have slightly different meanings.An iteration can correspond to a presentation of a single pair of images in the context of the invention, while an epoch corresponds to a complete passage through the Siamese system of all pairs of images.
[0099]
[0068] The concept of convergence (training and testing phases) is essential for validating the learning process. Convergence on pairs of training images is used to adjust the hyperparameters of the network model and to determine whether the model has converged, while convergence on pairs of test images is used to evaluate the final performance of the model on unknown images. The convergence of a neural network is the process by which the network learns to represent the relationships between inputs and outputs, and where the weights and biases of the neurons are adjusted to minimize prediction error.
[0100]
[0069] It is important to note that the convergence of a neural network system is not always guaranteed, and that it is possible that the network system may not converge towards an optimal solution.
[0101]
[0070] Training convergence occurs when the model's loss (or error) on the training image pairs decreases over the course of training iterations (epochs) and reaches a plateau. This means that the model has learned to represent the relationships between the inputs and outputs in the training pairs.
[0102]
[0071] Test convergence occurs when model loss on test pairs decreases over training iterations (epochs) and reaches a plateau. Model loss refers to pairs of images that were not recognized when they should have been. This means that the model generalizes well to new images that it did not see during training.
[0072] The goal is to achieve simultaneous convergence of the training and testing phases, meaning that the model learns to represent the relationships between inputs and outputs in the training images and generalizes well to new images.
[0103]
[0073] If convergence in the training phase is rapid, but convergence in the testing phase is slow or does not occur, this may indicate that the model is overfitted. That is, the model is too specialized for the training images and does not generalize well to new images.
[0104]
[0074] Conversely, if convergence in the test stage is rapid, but convergence in the training stage is slow or does not occur, this may indicate that the model is underfitted. This means that the model has not sufficiently learned the relationships between the inputs and outputs in the training images.
[0105]
[0075] In summary, the convergence of the training and testing stages is an important indicator of the performance of a CNN model, and it is essential within the scope of the invention to monitor these two metrics in order to adjust the hyperparameters and improve the performance of the CNN twin model implemented by the Siamese system. "Complete" learning includes both the training and testing phases; therefore, the Siamese system is only put into operation after undergoing the training phase, which includes both the training and testing stages.
[0106]
[0076] According to a first embodiment of the invention, the Siamese system is trained during the training phase, and more particularly during the training stage, with two balanced training sets: a first set of pairs of la bellized "recognized" image sets and a second set of pairs of la be Mized "unrecognized" image sets. Within the framework of the invention, the "training sets" are also referred to as "learning sets," and the terms "learning" and
[0107] "Training" are synonyms and used interchangeably.
[0108]
[0077] The sets are said to be balanced in that they comprise the same number, or at least a very close number, of image pairs. In the present case, each training set comprises at least 1000 pairs of image pairs and, preferably, more than 10,000 pairs of image pairs.
[0109]
[0078] Each pair in the first training set consists of two image tiles, each containing at least one texture with a random component. The two image tiles, when superimposed, are capable of exhibiting the Glass phenomenon. Figure 2 shows the two tiles or image tiles constituting a pair PI from the "recognized" set. As previously indicated by "tile tiles capable of generating the appearance of 'Glass' patterns," it should be understood that these are tiles from the superposition of which a human operator with normal visual acuity will certainly observe a Glass pattern, possibly after registration operations, such as manual registration. In Fig. 2, image tile Img 1 shows the result of this superposition, on which the Glass pattern is visible.In the first training set, the tiles that make up the same training image can, when simply superimposed without any other operation, either produce a Glass pattern or not produce a Glass pattern to an observer with normal visual acuity.
[0110]
[0079] Each training pair of the second set consists of two image tiles comprising at least one texture with a random component. The superposition of the two tiles cannot exhibit the Glass phenomenon, even after fine registration followed by the application of a residual transformation. Figure 2 shows two tiles or image tiles constituting a pair P2 of the set, which are "unrecognized." In Fig. 2, image tile Img 2 shows the result of superimposing the two image tiles of pair P2, on which no Glass pattern is visible.
[0111]
[0080] The Siamese twin system thus trained achieves a score of around 80% during the exploitation phases, without it being possible to significantly improve it even by increasing the size and number of training sets. It should be noted that in the first embodiment of the learning phase, no registration or residual transformation is performed on the images of the pairs in the recognized training set prior to their segmentation. However, it is possible to perform registration and / or residual transformation on the images of each recognized pair.
[0112]
[0081] According to a second embodiment of the invention, the same Siamese system as before is implemented, but its learning, training, is carried out with differently constituted training sets.
[0113]
[0082] Thus, according to this second embodiment, the Siamese system is trained with two balanced training sets: a first set of pairs of "Glass-direct" imagelets and a second set of pairs of "non-Glass" imagelets.
[0083] Each pair of "Glass-direct" imagelets in the first set consists of two imagelets comprising at least one texture with a random component. The two imagelets, after superposition, allow the Glass phenomenon to appear to an observer with normal visual acuity without the need for any prior processing or operation, either before or during the superposition. This is referred to as simple superposition.
[0114]
[0084] Each "non-Glass" image in the second set consists of two image tiles comprising at least one texture with a random component, the superposition of the two image tiles not being able to cause the Glass phenomenon to appear even after a fine registration operation followed by the application of a residual transformation.
[0115]
[0085] Once the learning has been carried out with the two sets "Glass direct" and "non-Glass", the CNN' neural network obtains a success rate of more than 90% in the exploitation phase.
[0116]
[0086] Training according to the second method results in a Siamese system that is less sensitive to image acquisition and registration conditions, making it more robust and facilitating its implementation in industrial and / or consumer applications for authenticating mass-produced physical objects. These improved performance characteristics are also found in the twin CNN1 and CNN2 networks that constitute the Siamese system thus trained.
[0117]
[0087] In the preceding example, two categories or classes of image pairs were identified: "Glass-direct" pairs and "non-Glass" pairs. It is possible to identify a third category of image pairs called "Glass-indirect." Each "Glass-indirect" pair consists of two images, each containing at least one texture with a random component. The simple superposition of these two images does not exhibit the Glass phenomenon to an observer with normal visual acuity. However, the superposition of the two images of a "Glass-indirect" pair exhibits a Glass pattern, either after coarse registration or after fine registration followed by a residual transformation. The image pairs of the specific categories "Glass-direct" and "Glass-indirect" both belong to the general category of image pairs "capable of generating the appearance of Glass patterns."
[0118]
[0088] In the context of this application, it can be said that the tiles of a "Glass Direct" image are normalized or normalized with respect to each other. This normalization can be carried out by an operator or semi-automatically with automatic registration and control by an operator of the overlay, or completely automatically with automatic registration and the automatic application of a predefined residual transformation.
[0119]
[0089] Surprisingly, the Siamese system trained according to the second embodiment proves capable of identifying certain "glass-indirect" pairs or imagelets as "recognized," meaning that it is able to identify non-normalized imagelets or pairs of imagelets even though the transformation from one to the other is outside the domain of residual transformation. This characteristic is particularly advantageous because it allows, in operation, the implementation of a coarse registration, which is faster and less resource-intensive than a fine registration.Similarly, the Siamese system trained according to the example of the second realization form is able to classify in the recognized category a pair of identical images which belong to the category "Glass-indirect" insofar as their simple superposition cannot make a Glass pattern appear but it is possible to make the Glass pattern appear by superimposing the two images one of which will have undergone a residual transformation.
[0120]
[0090] It should be noted that during training, the LOSS function must tend towards a zero value or at least a value as small as possible, minimal, for the "recognized" or "Glass-direct" sets.
[0121]
[0091] Once the Siamese system is trained, it is implemented to ensure, in a phase, called exploitation or execution, and in accordance with the invention, the authentication of a material subject by confronting, on the one hand, at least one image called reference of at least one authentication region of an authentic subject, the reference image comprising at least one texture with a random component, and on the other hand, a so-called candidate image of at least one authentication region of a candidate subject, the candidate image comprising at least one texture with a random component.
[0122]
[0092] Thus, for each authentication operation, at least one image acquired from a recognition region of a candidate material subject, called the candidate image, will be used, and at least one reference image acquired from a recognition region of a previously recorded reference material subject, called the reference image.
[0093] EP 3 380987 patent sets out acquisition conditions for the reference and candidate images suitable for their superposition to produce a Glass pattern. It should be noted in particular from this patent that the Glass pattern only appears in the case of an authentic subject and that there is a slight non-zero geometric transformation, referred to in the invention as a residual geometric transformation, between the verification (candidate) and authentication (reference) images acquired under the given conditions.In the theoretical case of a perfect superposition of strictly identical elements / images, no Glass pattern appears, even in the presence of an authentic subject. This underscores the necessity of this residual geometric transformation and the general advantage of implementing a relative movement or displacement, or even a deformation induced by a difference in shooting angle or viewpoint between the authentication and verification images. This property of Glass patterns provides significant robustness to the method according to the invention, as it is not necessary for the acquisition conditions of the candidate verification image to be strictly identical to the acquisition conditions of the reference authentication image. Thus, for example, the resolutions of the reference authentication and candidate verification images can be different.
[0123]
[0094] To make the best use of this property of Glass patterns, a training set can be implemented during the training phase, comprising direct Glass and / or indirect Glass type training pairs, each consisting of two distinct images / imagelets of the same texture with a random component with different shooting conditions, and whose superposition reveals or is likely to reveal a pattern analogous to a Glass pattern.
[0124]
[0095] During operation, the reference and candidate images are compared in a process implementing all or part of the Siamese system as described above, which has been trained as stated. By "compared," it is understood that the result of the process depends on these two images and that it is absolutely necessary to have them available in order to implement the invention. The reference and candidate images are two input data for the process according to the invention. It should be noted, however, that it is not strictly necessary to have the reference and candidate images available simultaneously, as will become clear later.
[0096] It should be noted that large original, candidate, and reference images can be used, and in that case, they must be cropped to fit the image tags that the twin networks CNN1 and CNN2 accept as input.For example, if the candidate and reference images are 3840 x 2880 pixels and each CNN twin network is configured to process 128 x 128 pixel images, the candidate and reference images can be divided into 660 images. To achieve this, and as shown in Figure 1, an input block (INPUT) is implemented before the CNN neural network. This block includes a slicing module (ECH) and successively provides each of the resulting images to the CNN1 and CNN2 twin networks.
[0125]
[0097] Fig.3 schematically illustrates the result of the processing carried out by the ECH slicing module. Thus, the ECH slicing module slices the reference image IMG ref and the candidate image IMG cand into tiles PAV in a common coordinate system so as to provide at the output of the slicing the set of pairs of homologous tiles; that is to say all the pairs each consisting of a tile of the reference image PAV ref and the tile of the candidate image PAV cand having exactly the same coordinates in the coordinate system as that of the tile PAV ref.
[0126]
[0098] During operation, the input block INPUT preferably includes, upstream of the ECH cutting module, a registration module 40 which ensures registration of the reference IMG images Ref and candidate Cand with respect to each other.
[0127]
[0099] Registration refers in particular to coarse registration, that is, an approximate alignment of images with each other, based on general characteristics such as size, orientation, or overall position. This can be carried out, for example, during the recognition stage by an operator equipped with a sensor (e.g., a smartphone), who manually positions the sensor against an authentication region of a physical object. Registration also refers, in particular, to fine registration, that is, a precise alignment of images, based on detailed characteristics such as contours, textures, or points of interest, and which can be down to the pixel or even sub-pixel level. It should be noted that in the context of coarse registration, the implementation of the residual transformation is not always necessary.It should also be noted that the concept of registration, within the scope of the invention, covers the registration of two images relative to each other, but also the registration of each of two images relative to the same third image, which could be called a pivot image. Thus, within the scope of the invention, the images to be processed can all be registered relative to the same pivot image. In the case of thumbnails, they are registered relative to the same pivot thumbnail.
[0128]
[0100] The INPUT block then includes a transformation module 41 which induces a residual transformation of the supplied images relative to each other. For this purpose, the transformation module 41 applies, for example, the residual transformation to one of the two images and not to the other. However, the transformation module 41 could operate differently insofar as the final result of the processing it performs corresponds to a residual transformation of the two supplied images relative to each other. For example, the residual transformation is a rotation of a few degrees combined with a rectilinear translation of about ten pixels. According to the example illustrated in Figure 1, the transformation module 41 is located between the registration module 40 and the clipping module ECH. However, it could be considered to place the transformation module 41 after the clipping module ECH.The transformation module 41 then works on the reference blocks Ref and candidate Cand before their provision to the twin networks.
[0129]
[0101] Fig. 4 shows the result of superimposing the images from a candidate image onto those from the reference image when the transformation module 41 is located between the registration module 40 and the ECH clipping module, while Fig. 5 shows the result of superimposing the images from the candidate image onto those from the reference image when the transformation module 41 is located after the ECH clipping module. In the context of Figs. 4 and 5, the candidate image and the reference image are images of the same region of the same material subject, which explains the presence of the Glass pattern in both cases.
[0130]
[0102] In the context of the invention and of the present application, the terms images, imagettes and tiles refer to objects, in the mathematical and computer science sense of the term, of the same nature, so that a treatment described in the context of the present application as being applied to one of these three types of object can, mutatis mutandis, be applied to the other two objects depending on the time when said treatment is implemented in the course of the process according to the invention.
[0131]
[0103] In the operational phase, the Siamese system does not include the LOSS stage and the two twin artificial neural networks operate in parallel, the first twin network CNN1 will process the imagettes img ref from the reference image IMG ref and the second twin network CNN2 will process the imagettes img cand from the candidate image IMG cand.
[0132]
[0104] Given the large number of processed images, it will be considered that unit recognition of the authentication region is achieved if a given percentage of the candidate images have been recognized as implementing the Glass phenomenon on the two original images considered. In this regard, it should be noted that authentication can be concluded if the Glass phenomenon, or the possibility of the Glass phenomenon, is identified for at least one presented candidate image.
[0133]
[0105] To perform the processing involved in this decision method, an OUTPUT block is implemented after the twin networks CNN1 and CNN2. This block collects the output of each twin network CNN1 and CNN2 for each of the images processed by them. The output of each twin network CNN corresponds to a descriptor of the corresponding image calculated by said twin network. The OUTPUT block performs the processing necessary for the authentication decision. This processing includes, at a minimum, a step to calculate the degree of similarity between the descriptor of the reference image and the descriptor of the candidate image, or between the descriptors of the reference images and the candidate images.
[0134]
[0106] In the latter case and according to a variant of the invention, the OUTPUT block can implement a statistical comparison with a similarity index between images, determine a given acceptance threshold, per pair of images or for a set of images located or not on the material subject, and issue an authentication decision possibly associating it with a confidence index of this decision.
[0135]
[0107] The validity of this decision can be audited and, if necessary, validated by a human user by presenting the latter with the superimposed reference and candidate images and allowing them to visually check for the presence or absence of a Glass pattern on this superposition.
[0136]
[0108] According to the example of the execution phase described above, the Siamese system comprises, for said execution phase, the two twin networks which are executed simultaneously or concurrently. However, this is not necessary for the implementation of the invention.
[0137]
[0109] Thus, according to another example of an execution phase according to the invention and as shown in Fig. 6, a single twin network of artificial neurons belonging to the trained Siamese set is implemented without the loss stage. This single twin network can be either of the twin networks CNN1 and CNN2 of the Siamese system from the training phase, insofar as after the latter they have the same characteristics or are even identical. For the purposes of this description, the single twin network is identified by the reference CNN. Furthermore, the twin network CNN is associated with the INPUT module, the operation of which differs from what has been described previously only in that it is adapted to deliver to the twin network CNN a single image to be processed or the imagelets of a single image to be processed, and insofar as the reference and candidate images are provided to it simultaneously, it would transmit them one after the other to the twin network CNN.
[0138]
[0110] The execution phase then includes a step of providing the twin network with the reference image IMGref so that it calculates a descriptor Dref of the reference image.
[0139]
[0111] The execution phase also includes a step of providing the twin network with the candidate image IMGcand so that it calculates a DCand descriptor of the reference image.
[0140]
[0112] The description phase then includes a step of calculating a degree of similarity between the Dref descriptor of the reference image and the Dcand descriptor of the candidate image. This calculation is performed by the OUTPUT block as described previously.
[0141]
[0113] In this example of the preferred form of the execution phase, the OUTPUT output block performs the same operations as described above to enable a decision regarding authentication.
[0142]
[0114] It should be noted that a descriptor calculated by a twin network trained according to the invention can be used as an index or an indexing tool for the image from which it is derived.
[0143]
[0115] In the context of the examples described above, the labels used to distinguish the training sets are "recognized", "not recognized", "Glass-direct", "Glass-indirect" and "Non-Glass" to facilitate the disclosure of the invention, but other labels could be used insofar as they allow the types of training sets to be distinguished from each other.
[0144]
[0116] Preferably, the image sets used in the training sets include a large number of image sets with random textures from different types of material. This configuration of the training sets results in a more efficient CNN neural network, as it is capable of identifying Glass patterns on a large number of superimposed image sets with random textures from a wide variety of materials. Furthermore, the CNN thus trained will be able to recognize a Glass pattern resulting from the superposition of random textures that were not present in the training sets.
[0145]
[0117] It should be noted that an artificial neural network exhibits, during operation, a behavior that reflects the learning stages it has undergone. Thus, an artificial neural network configured to be sensitive to Glass patterns in the context of images with random texture components will be recognizable among others, particularly by subjecting it to a set of tests containing such images and analyzing the responses provided by the artificial neural network under study.
[0146]
[0118] Within the framework of the invention, the labeling of training images or image sets can be carried out in different ways. Initially, the labeling of the training sets is performed by an operator. This is referred to as supervised training. However, after an initial supervised training phase, it is possible to perform a second, unsupervised automatic training phase in which the image labeling is carried out by a CNN neural network that has undergone initial supervised training.
[0147]
[0119] Similarly, the image sets belonging to the training sets can be derived from textured images with a random component of physical subjects, called natural image sets, or be image sets resulting from the superposition of synthetic images, called synthetic image sets. According to the invention, the training sets can comprise only natural image sets, only synthetic image sets, or a mixture of natural and synthetic image sets. Synthetic images are very well suited to simulating different lighting conditions (orientation, types, etc.), which can be very powerful for preparing a CNN for different shots of the candidate image and increasing its robustness.
[0148]
[0120] By way of example, the applicant advantageously used randomly generated Perlin texture images to create training sets for a CNN' neural network according to the invention, which, after training, proved effective in authenticating paper-type physical subjects.
[0121] Furthermore, the training sets can be created by one or more operators through manual or supervised semi-automatic operations. However, the creation of the training sets could be fully automated. Thus, training sets of "Glass-direct" imagelets or imagelet pairs can be automatically generated from series of continuously random textured images of various physical subjects.The generation of each "Glass-direct" image pair consists of extracting an image from a series of images, which forms the first image of the pair. The second image of the pair is formed by applying a residual transformation to the first image. For "Glass-indirect" pairs, the process is similar to that of "Glass-direct" pairs, except that instead of applying a residual transformation to the first image to create the second image, the first image is rotated between 90° and 270° to form the second image.
[0149]
[0122] Similarly, training sets of imagelets or imagelet pairs of "non-Glass" type can be automatically generated from series of images with a continuous random component texture of various material subjects. The generation of each pair of "non-Glass-direct" imagelets consists of extracting an imagelet from a textured image of two distinct subjects.
[0150]
[0123] Furthermore, within the scope of this application and the context of the invention, the terms blocks and modules may refer to purely software elements, purely hardware elements, or elements combining software and hardware implementations. This also applies to the implementation of the entire invention, which may be purely software-based, purely hardware-based, or a combination of dedicated software and hardware components.
[0151]
[0124] It should also be indicated that within the framework of the invention the input block can implement various types of computational processes to carry out the different processing it performs, that it can in particular implement specific artificial neural networks configured to perform the operations of cropping, cutting and applying transformations to the images to be processed.
[0152]
[0125] Furthermore, in the examples described above, the images to be processed are segmented before being provided to each neural network configured to be sensitive to Glass patterns. However, such a mode of operation is not strictly necessary for the realization of the invention, since it is possible to implement artificial neural networks sized for processing images from acquisitions without it being necessary to segment these images beforehand.
Claims
Demands 1. A method for authenticating a material subject consisting of comparing, on the one hand, at least one so-called reference image of at least one authentication region of an authentic subject, the reference image comprising at least one texture with a random component, and on the other hand, a so-called candidate image of at least one authentication region of a candidate subject, the candidate image comprising at least one texture with a random component, characterized in that it comprises: - in a training phase: - the implementation of a Siamese set comprising at least two twin artificial neural networks that are adapted to operate in parallel and whose outputs are connected during the learning phase to the same loss stage implementing a loss function, - a training step in which the Siamese set is trained with at least one training set comprising training pairs, at least some of which comprise two images whose superposition reveals or is likely to reveal a pattern analogous to a Glass pattern, the two images of the same training pair each being provided to a separate artificial neural network belonging to the Siamese set, - after the training phase, in an operational phase: - the implementation, without the loss stage, of a twin network of artificial neurons belonging to the Siamese set trained to be sensitive to Glass patterns in the context of images with random texture components, - a step of providing the twin network with the reference image so that the twin network can calculate a descriptor of the reference image, - a step of providing the twin network with the candidate image so that the twin network can calculate a descriptor of the candidate image, - a step of calculating a degree of similarity between the descriptor of the reference image and the descriptor of the candidate image.
2. Authentication method according to claim 1, characterized in that during the training phase, prior to their provision to the twin neural networks The images of each training pair are artificially calibrated against each other. 3 Authentication method according to claim 2, characterized in that during the training phase, after calibration and before their provision to the artificial neural networks, the images of each training pair are subjected to a residual transformation with respect to each other. 4 Authentication method according to any one of the preceding claims, characterized in that the training set comprises training pairs comprising two distinct images of the same texture with random component with different shooting conditions and whose superposition reveals or is likely to reveal a pattern analogous to a Glass pattern. 5 Authentication method according to any one of claims 1 to 4, characterized in that, during the exploitation phase, the descriptor of the reference image is recorded and the candidate image is provided to the twin network after this recording. 6 Authentication method according to any one of claims 1 to 5, characterized in that, during the exploitation phase, the candidate image is recalibrated against the reference image before being supplied to the artificial neural twin network. 7 Authentication method according to claim 6, characterized in that, after registration and before being supplied to the twin network of artificial neurons, the candidate image is subjected to a residual transformation with respect to the reference image. 8 Authentication method according to any one of claims 1 to 5, characterized in that, during the exploitation phase, the candidate image and the reference image are recalibrated against the same pivot image before being supplied to the twin artificial neural network. 9 Authentication method according to claim 8, characterized in that, after registration and before being supplied to the twin network of artificial neurons, the candidate image and / or the reference image is subjected to a residual transformation with respect to the pivot image.
10. Authentication method according to any one of the preceding claims, characterized in that each artificial neural network is adapted to process images or imagelets of given dimensions and when the images to be processed are larger, The images to be processed are cut into sub-images or imagelets of appropriate size which are successively submitted directly to each artificial neural network. 11 Authentication method according to claim 10 and claim 7 or 9, characterized in that the cutting takes place after the registration and before the processing by the artificial neural network. 12 Authentication method according to claim 10 and claim 7 characterized in that after cutting and before processing by the neural network each candidate image is subjected to a residual transformation with respect to the corresponding reference image. 13 Authentication method according to claim 10 and claim 7 characterized in that after cutting and before processing by the neural network each candidate image and / or reference image is subjected to a residual transformation with respect to the same pivot image or image. 14 Authentication method according to any one of the preceding claims, characterized in that after implementation of the artificial neural network, it includes a step of presenting to a user the superimposed reference and candidate images. 15 Authentication method according to any one of the preceding claims characterized in that each candidate image is taken from a video stream. 16 Authentication method according to any one of the preceding claims, characterized in that each artificial neural network is a convolutional neural network. 17 Product computer program comprising code instructions for performing a method according to any of the preceding claims of authenticating a material subject from a reference image and a candidate image, when the program is run on a computer. 18 A computer-readable storage means on which a computer program includes code instructions for executing a method according to any one of claims 1 to 17 for authenticating a material subject from a reference image and a candidate image. 19 Computer device comprising at least display means, image acquisition means, user input means, information storage means communicating with computing and control means configured to implement the method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Method for authenticating and / or checking the integrity of a subject
EP3380987A1
Method for matching a candidate image with a reference image
EP4396789A1
Method of augmented authentification of a material subject
US10990845B2
Non-counterfeitable document system
US4423415A
procedure D'AUTHENTIFICATION PAR MOTIF DE GLASS
FR3044451A3