Method for authenticating material subjects using a glass pattern by means of an artificial neural network

The use of a convolutional neural network for joint image processing of material subjects with random textures addresses the inefficiencies of existing authentication methods, providing efficient and robust authentication with reduced computational demands.

WO2026033103A1PCT designated stage Publication Date: 2026-02-12KERQUEST
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/072834
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-08
Filing Date
2025-08-07
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing authentication methods for material subjects require significant computing power and long computation times, which is impractical for real-time applications with limited resources and high energy consumption.

Method used

A method using an artificial neural network, specifically a convolutional neural network, to process joint images of authentication regions with random textures, enabling efficient authentication by recognizing Glass patterns formed from superimposed images, allowing for flexible acquisition conditions and reduced computational requirements.

Benefits of technology

Achieves reliable authentication with computation times suitable for portable devices, offering high flexibility and robustness against variations in acquisition conditions, while maintaining low energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025072834_12022026_PF_FP_ABST
    Figure EP2025072834_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for authenticating a material subject, the method consisting in comparing, on the one hand, at least one image, referred to as the reference image, of at least one authentication region for authenticating an authentic subject, the reference image comprising at least one random-component texture, and, on the other hand, at least one image, referred to as the candidate image, of at least one authentication region for authenticating a candidate subject, the candidate image comprising at least one random-component texture. According to the invention, the method implements at least one artificial neural network configured to be sensitive to Glass patterns in the context of random-component texture images.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR AUTHENTICATING MATERIAL SUBJECTS BY MEANS OF A NETWORK OF ARTIFICIAL NEURONS

[0002]

[0001] The present invention relates to the technical field of unitary authentication of material subjects with electronic means.

[0003]

[0002] In the above-mentioned field, it is known, notably from US patent 4423415, to identify physical objects by extracting a signature from a so-called authentication region comprising an essentially random three-dimensional intrinsic microstructure. Initially, a first extraction, of a so-called reference signature of an authentic object, is generally performed using electronic computing tools such as a computer, smartphone, or tablet, which implement complex pattern recognition algorithms. The reference signature is then recorded. Subsequently, a second extraction, of a so-called candidate signature of a candidate object, is performed using electronic computing tools similar to those used in the first step.The reference and candidate signatures are then compared with similarity or correlation comparison algorithms to determine their level of proximity and to deduce, if applicable, that the candidate subject is indeed the authentic subject.

[0004]

[0003] US patent 10990845 proposed a method for unit authentication of physical objects by calculating similarity vectors between, in particular, reference and candidate images of authentication regions of authentic and candidate physical objects. The methods and algorithms implemented in this patent require significant computing power for real-time implementation or long computation times with less power if high execution speed is not required. However, this constraint of significant computing power or long computation times proves to be a disadvantage when it is necessary to obtain computation times as short as possible while having limited computing power combined with low energy consumption.

[0005]

[0004] It therefore arose the need for alternative solutions which provide a solution to this problem and allow for reliable unit authentication of material subjects while using limited computing resources and contained energy consumption and which offer greater flexibility with regard to the conditions of acquisition of reference and candidate images than prior art processes implementing signature extraction.

[0006]

[0005] To achieve this objective, the invention relates to a method for authenticating a material subject consisting of comparing, on the one hand, at least one so-called reference image of at least one authentication region of an authentic subject, the reference image comprising at least one texture with a random component, and on the other hand, a so-called candidate image of at least one authentication region of a candidate subject, the candidate image comprising at least one texture with a random component. According to the invention, the method implements at least one artificial neural network, and the images are processed jointly by the same artificial neural network configured to be sensitive to Glass patterns in the context of images with a texture with a random component. Thus, according to the invention, the determination of authenticity is based on the result provided by the neural network(s) implemented.This result can, for example, be a numerical value, in which case a threshold can be used from which the authentication result is positive and below which there is no authentication.

[0007]

[0006] By joint processing of the reference and candidate images by the same artificial neural network, it should be understood that the two images are processed together by the same artificial neural network, as opposed to processing each image by a separate artificial neural network, or processing each image by the same neural network but sequentially, one image after another. In the context of the invention, the notion of joint processing implies that the two images to be processed are provided as input to the same artificial neural network, which will process them simultaneously. The images can then be provided as two separate inputs (e.g.channels) of the same artificial neural network or a single input of the same artificial neural network in the form of a combination of the two images to be processed, such as an image resulting from the superposition of the two images to be processed or an image resulting from the concatenation of the two images to be processed. This notion of joint processing also applies in the context of learning, also called training, of the artificial neural network; we then speak of training pairs for the pairs of training images that are processed jointly.

[0008]

[0007] According to one feature of the invention, in order to be sensitive to Glass patterns, the artificial neural network has been trained with at least one training set comprising training pairs, at least some of which comprise two images whose superposition reveals or is likely to reveal a pattern analogous to a Glass pattern.

[0009]

[0008] In a preferred embodiment, the invention implements at least one convolutional type artificial neural network.

[0010]

[0009] For the purposes of this invention, a convolutional artificial neural network, also called a convolutional artificial neural network, is a type of neural network comprising at least one, and preferably several, convolutional layers, and after the last convolutional layer, at least one layer of artificial neurons of another type, for example, at least one fully connected or densely connected layer of artificial neurons. Hereafter, the terms "convolutional artificial neural network," "convolutional neural network," and "convolutional neural network" shall be used interchangeably as synonyms. Similarly, within the scope of this invention, the terms "neural network" and "artificial neural network" are synonymous, and the term "neuron(s)" shall be understood to mean "artificial neuron(s)" unless otherwise specified.

[0011]

[0010] The implementation of convolutional artificial neural networks makes it possible to obtain satisfactory computation times with reasonable computing resources, such as those available in portable devices like smartphones or tablets, while exhibiting lower energy consumption than that required by prior art computational methods. Furthermore, the implementation of artificial neural networks allows for a certain flexibility in the acquisition conditions of reference and candidate images.

[0012]

[0011] For the purposes of the invention, a convolutional artificial neural network configured to be sensitive to Glass patterns in the context of images with random texture components is an artificial neural network that has been trained at least with images whose superposition reveals, or is likely to reveal, a pattern analogous to a Glass pattern, as described in EP 3,380,987 filed by the applicant, to which reference should be made. In the context of the invention, the artificial neural network implemented is configured to respond to images that, in combination, reveal or are likely to reveal said Glass pattern, or to an image resulting from the superposition of two images and revealing said Glass pattern.

[0013]

[0012] Thus, the present invention takes advantage of the applicant's demonstration that patterns similar to those obtained by Léon GLASS in articles in the journal NATURE vol.223 of August 9, 1969 pages 578 to 580 and Nature vol. 246 of December 7, 1973 pages 360 to 362, can appear by superimposing two images comprising respectively natural or artificial textures with random components resulting from the acquisition or even the photography at an adapted magnification or enlargement of the same intrinsic random multi-scale three-dimensional material structure of the same subject.The applicant has demonstrated that these Glass-type patterns appear only when continuous random textures originating from the same material structure and essentially residual geometric transformations of each other are superimposed. In practice, they do not appear when the continuous random textures are not sufficiently correlated or do not result from the acquisition of the same material structure corresponding to a subject's recognition region. Therefore, within the scope of the invention, Glass patterns enable unitary authentication, also known as unitary recognition. Conversely, if a Glass pattern is not observed, it is not possible to definitively conclude that the data is not authentic.Furthermore, the applicant highlighted the fact that, in the case of physical subject authentication, if said physical subject exhibits sufficient stability over time, images taken at different times, even if separated by several days, months, or years, can, through their superposition, generate such Glass patterns. Moreover, according to the invention, the authentic subject can undergo modifications after the authentication image is recorded while remaining authenticable, provided that a portion of the recognition region has not been significantly affected by these modifications, whether intentional or not.Furthermore, the applicant has demonstrated that the observation of Glass patterns is possible by superimposing two images of the same recognition region possessing a texture with a continuous random component, without the addition, extraction, or generation of discrete elements or discrete patterns as advocated in the prior art before patent EP 3380987 in the applicant's name. Such a Glass pattern appears only in the case of an authentic subject and only if there is a slight non-zero geometric transformation, referred to in the invention as a residual geometric transformation, between the candidate subject and the authentication image, or between the verification and authentication images acquired under the given conditions, the images having, where applicable, been previously registered with respect to each other.This property of Glass patterns provides the method according to the invention with great robustness, since it is not necessary for the acquisition conditions of the verification image to be strictly identical to the acquisition conditions of the authentication image. Thus, the resolutions of the authentication and verification images may differ. In the context of this application, an authentication image is also referred to as a reference image, while a verification image is also referred to as a candidate image.

[0014]

[0013] Within the framework of the invention, an authentication region, also called a recognition region, is a region that exhibits an intrinsic and random microstructure in that it results from the very nature of the authentication region of the material subject. In a preferred embodiment of the invention, each material subject used belongs to the families of subjects comprising at least one authentication region having an essentially random intrinsic structure that is not easily reproducible, that is to say, whose reproduction is difficult or even impossible in that it results in particular from a process that is not predictable at the observational scale.Such an authentication region with an essentially random, non-easily reproducible intrinsic continuous medium structure corresponds to Physical Unclonable Functions (PUFs), as defined in particular by Jorge Guajardo in the English-language Encyclopedia of Cryptography and Security, 01 / 2011 edition, pages 929 to 934. Preferably, the authentication region of a material subject conforming to the invention corresponds to an intrinsic nonclonable physical function, designated as "Intrinsic PUFs" in the aforementioned article.The applicant takes advantage of the fact that the random nature of the authentication region's microstructure is inherent or intrinsic to the very nature of the subject, resulting from its process of formation, development, or growth. Therefore, it is unnecessary to add a specific structure to the authentication region, such as an impression or engraving. Regarding the recognition region, according to the invention, its image possesses a texture that an observer with average visual acuity can perceive either with the naked eye or via optically and / or digitally enhanced zoom. Thus, small structures, or those perceived as small at the magnification, are seen as images containing a texture.

[0015]

[0014] In the context of the invention, the term texture refers to what is visible or observable in an image, while the term structure or microstructure refers to the material subject itself. The texture of interest in the recognition region is described as a texture with a random component in that it includes at least some randomness or irregularity. Thus, a texture or microtexture of the recognition region, called a texture-with-a-random component image, corresponds to an image of the structure or microstructure of the recognition region. Imaging such a recognition region under similar viewing conditions from nearby viewpoints provides images, each containing a texture with a random component that is the noisy reflection of its material structure.Such a texture with a random component inherits its unpredictability and independence from a texture with a random component from a completely different recognition region, from the element of randomness in the formation of their material structures.

[0016]

[0015] For the purposes of this invention, the term "image," whether candidate, reference, or training, means any type of image in the general sense of the term, and not solely an image comparable to a photograph. In other words, an image is not limited to an optical image resulting from the stimulation of the recognition region by visible light, but can instead be obtained by any type of physical stimulation, including but not limited to: ultrasound, far-infrared, terahertz, X-rays or gamma rays, X-ray or laser tomography, X-ray radiography, and magnetic resonance imaging. Thus, for the purposes of this invention, an image is, for example, the recording of the result of stimulation by any means whatsoever of a natural scene or a material subject. This recording can then be described as a natural image.This recording may be one-dimensional, corresponding, for example, to the recording of the variation over time or along a line of a single signal, or to the recording of the values ​​from a line of sensors. This recording may also be two-dimensional, as is the case with a photograph, which can be recorded in halftones, grayscale, or color. For the purposes of this invention, an "image" can therefore be a 1D signal, a 2D or 3D image, in grayscale or color, or an nD (n-dimensional) signal, for example, hyperspectral or RGB-D. For the purposes of this invention, an "image" can be the result of a single acquisition or extracted from a stream. Thus, within the scope of this invention, an "image" can be extracted from a video stream. In one embodiment, the image can be saved digitally.Furthermore, an image used in the process according to the invention can be a pre-segmented image, particularly when it comprises several physical subjects in the same scene or a repetition of the same motif. Of course, an image is not necessarily natural and can be synthetic, meaning it can be generated by a computer process with or without the assistance of a human operator. Within the scope of the invention, a natural image and a synthetic image have in common that they are in the same digital or analog recording format for processing within the same process. Optical and / or digital image enhancement pre-processing can also be applied to improve the image, for example, to achieve a better signal-to-noise ratio.

[0017]

[0016] According to a first embodiment of the invention, for their joint processing the candidate and reference images are superimposed and the artificial neural network is provided with an image resulting from the superposition of the candidate image and the reference image, the artificial neural network having been trained by means of superimposed random component texture images, the resulting images or superposition images having some of a Glass pattern and others not having a Glass pattern.

[0018]

[0017] Image superposition refers to any combination of all or part of at least two images to obtain a new image in which all or part of the original images contribute. According to a preferred feature of the invention, the candidate authentication, reference, and verification images do not undergo, for the purpose of superposition, any transformation during the verification phase other than operations to enhance or modify contrast, brightness, halftone transformation, color space changes such as conversion to grayscale or black and white, operations to modify saturation in certain hues, level inversion, or relative opacity modifications via the alpha channel.Thus, according to this preferred characteristic, images generally undergo so-called enhancement transformations that do not affect the ability to visually recognize the nature of the subject. Preferably, the applied transformations do not distort the images, particularly the continuous random textures they contain.

[0019]

[0018] In this context, superimposition consists in particular of:

[0020] - to superimpose one image on another by aligning the two images and stacking them on top of each other

[0021] - Calculate the difference between two images by subtracting the pixel values ​​of one image from the other / Weighted average: Calculate a weighted average of the two images - Blend two images using a blending function. This method combines the color channels of the two images to create a blending effect

[0022] - Perform a layering operation using "alpha-blending": This method involves using an alpha channel to control the transparency of the superimposed image.

[0023] - to perform a superposition with offset, namely to offset the superimposed image relative to the base image,

[0024] - perform a superimposition with a scaling modification, namely resize the superimposed image to compare areas of interest of different sizes,

[0025] - perform a superposition with variable opacity, namely adjust the opacity of the superimposed image to visually compare the images according to the desired opacity,

[0026] - to perform a mask overlay, that is, to create a mask to specify the areas to be overlaid in each image. It should be noted that a Glass pattern may appear when viewing the overlay of more than two images. Therefore, within the scope of this invention, the term "overlay" refers to the superposition of at least two images.

[0027]

[0019] According to a second embodiment of the invention, the candidate image and the reference image are jointly provided to the artificial neural network without any overlap. The artificial neural network has been trained using pairs of non-overlapping images with random textures, called training pairs. For some of these pairs, the overlap of these images reveals, or is likely to reveal, a Glass pattern, while for others, the overlap does not reveal, and is not likely to reveal, a Glass pattern.

[0020] According to a variant of this second embodiment, for their joint processing, the two images to be processed—either from each training pair or the reference and candidate images—are combined by juxtaposing them without overlap to form a single image, which is then provided to the artificial neural network.Preferably, this assembly is performed after any necessary realignment, possibly followed by a residual transformation. Within the framework of the invention, a juxtaposition assembly without overlap corresponds to a concatenation.

[0028]

[0021] According to another variant of this second embodiment, for their joint processing the two images to be processed are provided simultaneously to the same artificial neural network and / or to two distinct inputs of this same artificial neural network. The term "input" is here synonymous with "channel".

[0029]

[0022] In this second embodiment, and according to a variant of the corresponding learning process, learning is carried out using, on the one hand, images whose superposition is known to generate a Glass pattern, without any particular precautions regarding their presentation to the artificial neural network to be trained, and, on the other hand, images whose superposition is known not to generate Glass patterns as defined in the invention. By "no particular precautions," it should be understood that the images of the same pair have not been subjected to any registration, whether during the training or processing phase.

[0030]

[0023] By proceeding in this way, an artificial neural network configured to be sensitive to the Glass pattern was obtained, achieving an identification rate of approximately 80% during operation, without it being possible to substantially improve this authentication rate by increasing the size of the training datasets. By "images whose superposition is known to be likely to generate Glass patterns," it should be understood that these are images from the superposition of which a human operator will certainly observe a Glass pattern, possibly after registration operations have been performed, for example, manually.

[0031]

[0024] Such an identification rate of 80% may be satisfactory in certain applications, particularly when it is possible to have a second check performed by an operator implementing the invention as described in patent EP 3380987 on behalf of the applicant. However, this rate is insufficient when it is necessary to process a large number of physical subjects, or their images, within a limited and short timeframe, thus restricting the possibility of human intervention.

[0032]

[0025] The need therefore arose to improve the authentication or recognition rate. To this end, a variant of the second embodiment of the method according to the invention proposes to register the images of each training pair relative to each other and then apply a residual transformation to them prior to their submission to the artificial neural network during the training phase. It was found that, quite surprisingly, implementing registration and then applying a residual transformation prior to presenting the images of the training pairs yielded better results than when no precautions were taken. By proceeding in this way, it was possible to achieve authentication rates of around 95% or even higher during the operational phase.

[0033]

[0026] For the purposes of this invention, a residual transformation or residual geometric transformation is a slight, non-zero geometric transformation to be applied locally to the chosen image or images, whether rigid or not, linear or non-linear, at at least one fixed or quasi-fixed point. Among the applicable geometric transformations, it is thus possible to implement the transformations described by Leon Glass in his 1973 and 2002 articles cited above. A quasi-fixed point is understood to be a point that, after the residual geometric transformation, undergoes a displacement of small magnitude compared to the maximum displacement caused by the residual geometric transformation.In the theoretical case of a perfect superposition of strictly identical elements / images, no Glass pattern appears even in the presence of an authentic subject. Hence the necessity of this residual geometric transformation and the general advantage of implementing a relative movement or displacement, or even a deformation induced by a difference in the shooting angle or viewpoint between the acquisitions of the authentication image and the verification image.

[0027] It should be noted that, according to one embodiment of the invention, the training pairs are calibrated, and all or some do not undergo a residual transformation before their joint submission to the artificial neural network during the training phase.

[0034]

[0028] According to one embodiment of the invention, prior to their joint processing, the candidate and reference images are registered with respect to each other. In the first embodiment of the invention, the registration takes place before the superposition is provided to the artificial neural network for joint processing.

[0035]

[0029] According to the invention, the registration of two images can be performed either relative to each other or each relative to a common registration image distinct from the two images being registered. Thus, the images of a training pair can each be registered relative to a registration image that will be used for the registration of the images of all the training pairs. This registration image, which can be called the pivot image, can also be used during the processing phase for the registration of the candidate and reference images.

[0036]

[0030] In this variant and according to a feature of the invention, after registration and before their joint processing, the candidate and reference images undergo a residual transformation with respect to each other. In the first embodiment of the invention, the residual transformation is applied after registration and before the superposition is provided to the artificial neural network for joint processing.

[0037]

[0031] It should be noted that these residual recalibrations and transformations can be applied to the constitutive images of the training pairs, as presented above.

[0038]

[0032] According to one feature of the invention, the method according to the invention in its first embodiment is used to label learning pairs used for learning, also called training, the method according to the invention in its second embodiment.

[0039]

[0033] According to an embodiment of the method according to the invention, the artificial neural network is adapted to process images or imagelets of given dimensions, and when the images to be processed are larger, they are divided into sub-images or imagelets of a suitable size, which are then submitted directly to the artificial neural network.

[0034] In the context of this application, and when reference is made to the processing performed by the invention, the term "image" is preferably used for the result of the acquisition, while the term "imagelet" is used for the object that is processed by each neural network. However, it should be noted that images and imagelets are of the same nature, differing only in size or dimensions, so that these terms may be used interchangeably depending on the context, without hindering the understanding of the invention by a person skilled in the art.Indeed, the terms image(s) and thumbnail(s) are used to facilitate understanding of the description of the invention and should therefore not be interpreted as limiting. Similarly, in some cases, sub-images may also be referred to as a block of text; this term, also used to facilitate the description of the invention, should not be interpreted as limiting.

[0040]

[0035] According to a feature of this method of implementing the process according to the invention, the candidate and reference images are recalibrated prior to their cutting.

[0041]

[0036] According to a preferred feature of the invention, the registration is carried out before the cutting while the residual transformation is applied after the cutting.

[0042]

[0037] According to a characteristic of the method according to the invention, each candidate image is derived from a video stream. This characteristic makes it possible, in certain implementation contexts, to avoid the residual registration and transformation operations applied to the candidate and reference images prior to their processing by each neural network. This is particularly the case when the video stream originates from an acquisition of a physical subject to be authenticated, performed by an operator who carries out a crude registration during the acquisition process.

[0043]

[0038] According to a preferred feature of the invention, at least one artificial neural network implemented is a convolutional neural network.

[0044]

[0039] The invention also relates to a computer program product comprising code instructions for executing a method according to the invention for the one-time authentication of a physical object from a reference image and a candidate image, when the program is executed on a computer. Such a computer program product can then be embedded or stored in a computer or similar device. For the purposes of this invention, the term "computer" should be understood in a broad sense as including, in particular, a personal computer, a smartphone, a tablet, a virtual or physical server, or any computing and processing unit capable of implementing a computer program and adapted for carrying out the invention.

[0045]

[0040] The invention also relates to a means of storage readable on a computer equipment on which a computer program includes code instructions for the execution of a method according to the invention of authenticating a material subject from a reference image and a candidate image.

[0046]

[0041] The invention also relates to a computer device comprising at least display means, image acquisition means, user input means, information storage means communicating with computing and control means configured to implement the method according to the invention of authenticating a material subject from a referenced image of a candidate image.

[0047]

[0042] The different modes of implementation, forms of embodiment, characteristics and variants of the invention can be implemented with each other in different combinations insofar as they are not mutually exclusive or incompatible with each other.

[0048]

[0043] Various other features and variants of the invention will become apparent from the description below, made in relation to the figures in which: Fig. 1 is a schematic representation of a first embodiment of the invention implementing a convolutional neural network; Fig. 2 illustrates the superposition of imagelets resulting from the processing of a reference image and a candidate image, in which the clipping is carried out after registration of the two images and application of a residual transformation; Fig. 3 illustrates the superposition of imagelets resulting from another processing of the reference image and the candidate image used in Figure 2, in which the images are registered and then clipped, the residual transformation being applied to the imagelets obtained before their superposition; Fig.Figure 4 shows image sets of training sets used for training convolutional neural networks implemented within the framework of the invention, the right part of the figure showing the result of superimposing the image sets located on the same line on the left, Figure 5 is a schematic representation of a second embodiment of the invention, Figure 6 is a schematic representation of the operation of an image slicing module implemented within the framework of the second embodiment of the invention.

[0049]

[0044] In the figures, the elements common to the different embodiments bear the same reference numerals. Furthermore, the different embodiments, presented in relation to the figures, correspond to non-limiting examples of possible executions and implementations of the invention.

[0050]

[0045] As previously stated, the invention implements a convolutional artificial neural network to ensure the authentication of physical objects from images of an authentication region of these objects, the images comprising at least one texture with a random component. In a preferred embodiment, the invention enables unitary authentication, that is, a reference image, also called an authentication image, corresponds to one and only one physical object.

[0051]

[0046] Thus, unit authentication means unit recognition of a region of a material subject. This recognition can have a higher or lower probative value depending on the criticality of the use case considered and the measures implemented to increase this probative value, such as, but not limited to: the involvement or not of a trusted third party, the selection of highly sophisticated acquisition sensors in terms of resolution and illumination conditions for acquiring reference images and the use of the same sensors for acquiring candidate images, and complete control of the IT environment used, without this list being exhaustive or limiting.

[0052]

[0047] According to the invention, the authentic subject can undergo modifications after the recording of the reference image, also called the authentication image in the context of patent EP 3380987, while remaining authenticable to the extent that a part of the authentication region has not been profoundly affected by these voluntary or involuntary modifications.

[0053]

[0048] In order to be implemented on portable devices and / or to avoid requiring significant computing resources, which are also energy-intensive, the invention proposes implementing artificial neural networks that could be described as frugal in that they comprise a limited or even reduced number of layers. In a preferred embodiment, the invention implements convolutional neural networks comprising a limited number of convolutional neuron layers and fully or densely connected neuron layers.

[0054]

[0049] Among the examples of artificial convolutional neural networks, also designated by the abbreviation "CNN" (Convolutional Neural Network), that can be implemented by the invention, two examples can be cited by way of example: the LetNet or LetNeT-5 neural networks and the AlexNet neural networks. It is also possible to refer to the French and English Wikipedia pages entitled "convolutional neural network" and "convolutional neural network," respectively, for other examples of convolutional neural networks and explanations of their structure.

[0055]

[0050] According to a preferred embodiment of the invention, a convolutional artificial neural network is used, comprising a stack of processing layers, namely:

[0056] - Convolutional layers (CONV) that process data from a receiving field, namely an image,

[0057] - POOL pooling layers that allow information to be compressed by reducing the size of the intermediate image

[0058] - correction layers, often incorrectly called ReLU (rectified linear activation function), - completely or densely connected layers, FDC

[0059] - LOSS layer (loss function) which can also be called loss layer.

[0060]

[0051] It should be noted that for some authors the LOSS layer is considered not to be part of the neural network; therefore, within the scope of the invention, a neural network does not necessarily have such a layer. Furthermore, the so-called LOSS layer is present only during the learning phase and not during the execution phase.

[0061]

[0052] Thus, and as can be seen from Figure 1, an example of such a convolutional artificial neural network, designated collectively as CNN, is configured to process a rectangular image img of mxn pixels, preferably with m and n greater than or equal to 64, for example, 128x128 pixels, it being understood that m and n are not necessarily equal. The CNN neural network comprises a convolution processing block 1 followed by a processing block of fully or densely connected artificial neurons 2.

[0062]

[0053] According to the illustrated example, the CNN network further comprises, between the convolution block 1 and the processing block 2 with completely or densely connected artificial neurons, a FLAT layer for processing the output of the convolution block 1 before supplying it to the block 2.

[0063]

[0054] According to the illustrated example, the convolution block 1 comprises successively and in this order:

[0064] - a first convolution layer 11,

[0065] - a first correction layer 12,

[0066] - a first layer of pooling 13,

[0067] - a second convolution layer 14,

[0068] - a second correction layer 15,

[0069] - a second layer of pooling 16,

[0070] - a third convolution layer 17

[0071] - a correction layer 18 which happens to be the last layer of convolution block 1.

[0072]

[0055] The first convolution layer 11 is, in this case, a 2D convolution layer parameterized to process the entire image using 3x3 pixel tiles (kernels or filters) with a step of 1 pixel. The first correction layer 12 implements a Rectified Linear Unit (ReLU) activation function to process the result of the first convolution layer 11. The first pooling layer 13, as illustrated in the example, performs maximum pooling for 2D spatial data with a 2x2 window and a step (Stride) of 2, on the result of the processing performed by the first correction layer 12.

[0073]

[0056] The second convolution layer 14 is, in this case, a 2D convolution layer parameterized to process the output of the first pooling layer 13 by 5x5 pixel tiles with a step of 2 pixels. The second correction layer 15 implements a Rectified Linear Unit (ReLU) activation function to process the result of the second convolution layer 14. The second pooling layer 16 ensures, according to the illustrated example, a maximum pooling operation for 2D spatial data with a 2x2 window and a step of 2 on the result from the second correction layer 15.

[0074]

[0057] The third convolution layer 17 is, in this case, a 2D convolution layer parameterized to process the result of the second pooling layer 16 by 5x5 pixel tiles with a step of 2 pixels. The third correction layer 18 here implements a Rectified Linear Unit (ReLU) activation function to process the result of the third convolution layer 17. In the illustrated example, the third correction layer 18 is the last layer of the convolution block 1.

[0075]

[0058] An example of code for defining convolution block 1 as described above in PyTorch is as follows: self.cnnl = nn.Sequential nn.Conv2d(l, 128, kernel_size=3, stride=l), nn.ReLU(inplace=True), nn.MaxPool2d(2, stride=2), nn.Conv2d(128, 256, kernel_size=5, stride=2), nn.ReLU(inplace=True), nn.MaxPool2d(2, stride=2), nn.Conv2d(256, 512, kernel_size=5, stride=2), nn.ReLU(inplace=True),

[0076]

[0059] The CNN neural network includes at the output of the convolution block 1 a FLAT processing layer, having the reference 19, which ensures the flattening in the form of a single-row vector or matrix of the result of the processing from the third correction layer 18.

[0077]

[0060] Downstream of the processing layer 19, the CNN neural network comprises the processing block 2, which is a fully or densely connected neural network. According to the illustrated example, block 2 comprises, successively and in this order:

[0078] - a first layer of 20 densely or completely connected neurons

[0079] - a first layer of correction 21,

[0080] - a second layer 22 of densely or completely connected neurons,

[0081] - a second correction layer 23,

[0082] - a third layer 24 of densely or fully connected neurons which, according to the illustrated example, is the last layer of block 2.

[0061] According to the illustrated example, the first layer 20 is a linear layer of artificial neurons receiving 512 inputs and delivering 1024 outputs. By linear layer, it is understood that a line of neurons is a single layer of neurons.

[0083]

[0062] The first correction layer 21, of block 2, here implements an activation function of type ReLU for in English "Rectified Linear Unit" applied to each of the 1024 outputs of layer 20.

[0084]

[0063] The second layer 22 is a linear layer of artificial neurons receiving 1024 inputs and delivering 256 outputs.

[0085]

[0064] The second correction layer 23, of block 2, here implements an activation function of type ReLU for in English "Rectified Linear Unit" applied to each of the 256 outputs of layer 22.

[0086]

[0065] Finally, the third and last layer 24 is a linear layer of artificial neurons receiving 256 inputs and delivering 1 output.

[0087]

[0066] An example of code for the definition of block 2, for a fully or densely connected neural network, as described above in PyTorch language is as follows: self.fcl = nn.Sequential nn.Linear(512, 1024), nn.ReLU(inplace=True), nn.Linear(1024, 256), nn.ReLU(inplace=True), nn.Linear(256, 1),

[0088]

[0067] According to the illustrated example, the CNN neural network finally includes an output layer 24 which, in this case, ensures a normalization of the value delivered by the third and last layer 24 of block 2, in the form of a real number, between 0 and 1.

[0089]

[0068] It should be noted that the different values ​​of the functions mentioned in the Pytorch code examples are designated by the generic term hyperparameters; these are data which are not automatically updated during the learning phases, also called training phases, as opposed to the parameters which are, such as, for example, the weight of the connections between artificial neurons and the biases of these.

[0090]

[0069] According to the illustrated example, an input block (INPUT) is implemented upstream of the CNN. This block performs various processing operations on each image or pair of images to be processed before they are supplied to the CNN neural network. For example, when the image to be processed (IMG) is larger than the image (img) that can be processed by the CNN neural network, the input block (INPUT) will perform a segmentation or decomposition of the image (IMG) into sub-images or imagelets (img) that will be directly supplied to the CNN neural network. In other words, all processing operations that can be performed by the input block (INPUT) will be carried out prior to this segmentation. In the context of this application, the segmentation performed by the segmentation module corresponds to segmenting an image into a multitude of smaller imagelets or into a multitude of smaller blocks.

[0091]

[0070] The training, also called learning, of an artificial neural network takes place according to a number of phases well known to those skilled in the art, which are as follows:

[0092] 1 / Initialization of weights: The weights of the connections between neurons are initialized with small random values.

[0093] 2 / Presentation of training images: The training images are introduced into the network, starting with the input layer and propagating towards the output layer of the network.

[0094] 3 / Output calculation: The network calculates the output for the given input, by applying non-linear activation functions in each neuron.

[0095] 4 / Error calculation: The error between the actual output and the predicted output is calculated, usually using a loss function.

[0096] 5 / Backpropagation of error: The error is propagated back through the network, through each layer, by adjusting the weights of the connections between neurons.

[0097] 6 / Updating the weights: The weights are adjusted according to the error gradient, using an optimization method such as gradient descent.

[0098] 7 / Repetition: Steps 2 to 6 are repeated for a large number of epochs (iterations), until the network achieves acceptable accuracy on the training images.

[0099] 8 / Evaluation / Test: The trained network is evaluated on a set of test images to estimate its performance.

[0100]

[0071] An "epoch" in the context of machine learning corresponds to a complete pass through the CNN of all the training (respectively, test) images or pairs of images. In other words, an epoch is a complete iteration in which the model sees all the training (respectively, test) images or pairs of images once, and the weights and biases are updated accordingly. For example, if the set of training (respectively, test) images or pairs of images contains 1000 images or pairs of images, an epoch corresponds to the presentation of these 1000 images or pairs of images to the CNN, and the updating of the weights and biases after each batch of image pair(s). It is important to note that the term "epoch" is often used interchangeably with the term "iteration," but they have slightly different meanings.An iteration can correspond to a presentation of a single image or pair of images in the context of the invention, while an epoch corresponds to a complete pass through the CNN of all the images or pairs of images.

[0101]

[0072] The concept of convergence (training and testing phase, also referred to as training and validation phase) is essential for validating the learning process. Convergence on training images is used to adjust the hyperparameters of the network model and to determine whether the model has converged, while convergence on test images is used to evaluate the final performance of the model on unknown images. The convergence of a neural network is the process by which the network learns to represent the relationships between inputs and outputs, and where the weights and biases of the neurons are adjusted to minimize prediction error.

[0102]

[0073] It is important to note that the convergence of a neural network is not always guaranteed, and that it is possible that the network may not converge towards an optimal solution.

[0103]

[0074] Training convergence occurs when the model's loss (or error) on the training images or pairs of images decreases over the course of training iterations (epochs) and reaches a plateau. This means that the model has learned to represent the relationships between the inputs and outputs in the training images.

[0104]

[0075] Convergence of the test occurs when model loss on the test images decreases over the course of training iterations (epochs) and reaches a plateau. Model loss refers to the images or pairs of images that were not recognized when they should have been. This means that the model generalizes well to new images that it did not see during training.

[0105]

[0076] The objective is to achieve simultaneous convergence of the training and testing phases, which means that the model learns to represent the relationships between inputs and outputs in the training images and generalizes well to new images.

[0106]

[0077] If convergence in the training phase is rapid, but convergence in the testing phase is slow or does not occur, this may indicate that the model is overfitted. That is, the model is too specialized for the training images and does not generalize well to new images.

[0107]

[0078] Conversely, if convergence in the test phase is rapid, but convergence in the training phase is slow or does not occur, this may indicate that the model is underfitted. This means that the model has not sufficiently learned the relationships between the inputs and outputs in the training images.

[0108]

[0079] In summary, the convergence of the training and testing phases is an important indicator of the performance of a CNN model, and it is essential to monitor these two metrics to adjust the hyperparameters and improve the performance of the implemented CNN model. "Full" learning includes both the training and testing phases; therefore, a neural network is only put into operation after undergoing both training and testing.

[0109]

[0080] Once the CNN neural network has been trained, it is implemented to ensure, in a phase, called exploitation or execution, and in accordance with the invention, the authentication of a material subject by confronting, on the one hand, at least one so-called reference image of at least one authentication region of an authentic subject, the reference image comprising at least one texture with a random component, and on the other hand, a so-called candidate image of at least one authentication region of a candidate subject, the candidate image comprising at least one texture with a random component.

[0110]

[0081] Thus, for each authentication operation, at least one image acquired from a recognition region of a candidate material subject, called the candidate image, will be used, and at least one reference image acquired from a recognition region of a previously recorded reference material subject, called the reference image.

[0111]

[0082] EP 3 380987 patent sets out acquisition conditions for the reference and candidate images suitable for their superposition to reveal a Glass pattern. It should be noted in particular from this patent that the Glass pattern appears only in the case of an authentic subject and that there is a slight non-zero geometric transformation, referred to in the invention as a residual geometric transformation, between the verification, candidate, and authentication, reference images acquired under the given conditions.In the theoretical case of a perfect superposition of strictly identical elements / images, no Glass pattern appears, even in the presence of an authentic subject. This underscores the necessity of this residual geometric transformation and the general advantage of implementing a relative movement or displacement, or even a deformation induced by a difference in shooting angle or viewpoint between the authentication and verification images. This property of Glass patterns provides significant robustness to the method according to the invention, as it is not necessary for the acquisition conditions of the candidate verification image to be strictly identical to the acquisition conditions of the reference authentication image. Thus, for example, the resolutions of the reference authentication and candidate verification images can be different.

[0112]

[0083] In the case of the method according to the invention, the reference and candidate images are compared in a process implementing a neural network such as, for example, described above and which has been trained as stated. By "compared" it should be understood that the result of the process depends on these two images and that it is absolutely necessary to have them in order to implement the invention, and that the reference and candidate images are two input data for the method according to the invention which are processed jointly.

[0113]

[0084] Note that large original, candidate, and reference images can be used, but they must then be cropped to fit the imagelets that the chosen CNN accepts as input. For example, if the candidate and reference images are 3840 x 2880 pixels and the CNN is configured to process 128 x 128 pixel imagelets, the candidate and reference images can be cropped into 660 imagelets. To this end, and as shown in Figure 1, an input block (INPUT) is implemented before the CNN. This block includes a cropping module (ECH) and successively provides each of the resulting cropped imagelets to the CNN.

[0114]

[0085] Given the large number of processed images, it will be considered that unit recognition of the authentication region is achieved if a given percentage of the candidate images has been recognized as exhibiting the Glass phenomenon on the two original images considered. In this regard, it should be noted that authentication can be concluded if the Glass phenomenon is identified for at least one presented candidate image. In order to perform the processing involved in this decision method, an OUTPUT block is implemented after the CNN network. This block collects the CNN network output for each of the images processed by the CNN and performs the processing necessary to determine whether or not images exhibit a Glass phenomenon, or are likely to exhibit one in the case of superimposition with or without registration.This decision-making process is automated and can be validated by an operator viewing the result of the overlay of the collected images or pictures.

[0115]

[0086] According to one variant, the OUTPUT block can implement a statistical comparison with a similarity index between images, determine a given acceptance threshold, per pair of images or for a set of images located or not on the material subject, and issue an acceptance decision possibly associating it with a confidence index of this decision.

[0116]

[0087] In a first embodiment of the invention, more particularly illustrated in Figure 1, the reference image IMG Ref and the candidate image IMG Cand are submitted to the CNN neural network for joint processing after being superimposed one on top of the other. For this purpose, the input block INPUT comprises a first module 30 which first performs registration of the reference and candidate images and then induces a residual transformation of one of these two images with respect to the other.

[0117]

[0088] Registration refers in particular to coarse registration, that is, an approximate alignment of images with each other, based on general characteristics such as size, orientation, or overall position. This can be carried out, for example, during the recognition stage by an operator equipped with a sensor (e.g., a smartphone), who manually positions the sensor against an authentication region of a physical object. Registration also refers, in particular, to fine registration, that is, a precise alignment of images, based on detailed characteristics such as contours, textures, or points of interest, and which can be down to the pixel or even sub-pixel level. It should be noted that in the context of coarse registration, the implementation of the residual transformation is not always necessary.It should also be noted that the concept of registration, within the scope of the invention, covers the registration of two images relative to each other, but also the registration of each of two images relative to the same third image, which could be called the pivot image or registration image. Thus, within the scope of the invention, the images to be processed can all be registered relative to the same pivot image.

[0089] Following the first module 30, the input module INPUT includes a second module 31 which performs a superposition of the reference and candidate images after their processing by the first module 30. The superposition image from the second module 31 is then provided to the slicing module ECH. According to this example, the slicing occurs after the preliminary operations performed on the reference and candidate images, and just before the resulting imagelets (img) are provided to the neural network CNN.However, this order is not strictly necessary, and other sequences of the registration, residual transformation, and splitting steps can be considered. For example, the splitting could occur after registration but before the residual transformation, which would then be applied to the pairs resulting from the splitting.

[0118]

[0090] Figure 2 illustrates the result of the clipping of the superposition of two images which have undergone registration and residual transformation before their clipping, while Figure 3 illustrates the result of the superposition of the imagelets from the same two images which have undergone registration then clipping and finally a residual transformation applied to the imagelets.

[0119]

[0091] According to this first embodiment, the neural network is trained with two balanced training sets: a set of "recognized" labeled images consisting of images resulting from the superposition of images comprising at least one texture with a random component, this superposition exhibiting the Glass phenomenon, and a set of "unrecognized" labeled images consisting of images resulting from the superposition of images comprising at least one texture with a random component, this superposition not exhibiting a Glass phenomenon. The sets are said to be balanced in that they comprise the same number, or at least a very close number, of images. In the present case, each training set comprises at least 1000 images and, preferably, more than 10,000 images.

[0092] Figure 4 shows an example of an image belonging to the set of recognized labeled images and exhibiting a Glass pattern.This image, imgl, is the result of superimposing the two images located to its left on the same line, which form a pair, P1, of training images, labeled "recognized" or "Glass." Figure 4 also shows an example of an image belonging to the set of images labeled "unrecognized" and not exhibiting a Glass pattern. This latter image is the result of superimposing the two images located to its left on the same line, which form a pair, P2, of training images, labeled "unrecognized" or "Non-Glass." The training set includes "Glass" training pairs and "non-Glass" training pairs.

[0120]

[0093] In a second embodiment more particularly illustrated in Figure 5, the candidate and reference images are processed jointly without any superimposition of the latter prior to their provision to the CNN neural network.

[0121]

[0094] The CNN' neural network implemented in this second embodiment is of a model analogous to the CNN described in relation to Figure 1. It differs in that the image img' processed by the CNN' network will be twice the size of the image img processed by the CNN network. Thus, in this second embodiment, the image img' will have a size of m' x n' pixels, with, for example, one of the two m' or n' greater than 128 pixels and the other greater than 64 pixels. In this case, the image will have dimensions of 128 by 256 pixels.

[0122]

[0095] The reference and candidate images will generally have a larger dimension than the imagelet img', so the ECH slicing module of the input block is configured to slice the candidate image and the reference image into PAV tiles, each of which has a size and shape corresponding exactly to half of the imagelet img'. As shown in Figure 5, the ECH slicing module slices the reference image IMG ref and the candidate image IMG cand into PAV tiles in a common coordinate system so as to provide at the output of the slicing the set of pairs of corresponding tiles; that is to say, all the pairs, each consisting of a tile of the reference image PAV ref and a tile of the candidate image PAV cand having exactly the same coordinates in the coordinate system as those of the tile PAV ref.

[0123]

[0096] According to the illustrated example, the tiles of each pair will be concatenated, juxtaposed, to constitute each image Img' provided to the neural network CNN'. This operation will be performed, according to the illustrated example, by the slicing module but could also be performed by another module separate from the slicing module. In the context of the invention and the present application, the term "pair" refers either to the result of the concatenation of two tiles, two images, or two images, or to the two tiles, two images, or two images in that they are intended to be processed jointly without having been assembled or concatenated.

[0097] In the second embodiment, the OUTPUT block will have the same mode of operation as that of the embodiment described in relation to Figure 1.

[0098] According to a first variant of this second embodiment, the neural network is trained with two balanced training sets: a first set of images labeled "recognized" and a second set of images labeled "not recognized". Each image in the first set is formed by the concatenation of two image tiles comprising at least one texture with a random component, the two tiles being capable, after superposition, of exhibiting the Glass phenomenon. The PI pair in Figure 4 shows two tiles which, when concatenated, form an image from the "recognized" set.As previously indicated by "tiles capable of producing Glass patterns," this refers to tiles whose superposition will reliably reveal a Glass pattern to a human operator with normal or average visual acuity, possibly after some registration, such as manual registration. In the first training set, the tiles comprising a single training image can, upon simple superposition without any further operation, either produce a Glass pattern or not produce a Glass pattern to an observer with normal visual acuity.

[0124]

[0099] Each image in the second set is formed by concatenating two image tiles, each containing at least one texture with a random component. The superposition of the two tiles does not exhibit the Glass phenomenon, even after fine registration followed by the application of a residual transformation. Pair P2 in Figure 4 shows two tiles which, when concatenated, form an image from the "unrecognized" set.

[0125]

[0100] The CNN' neural network, thus trained, achieves a score of around 80% during the operational phases, with no significant improvement possible even by increasing the size and number of training sets. It should be noted that, in the first variant of the second embodiment of the invention, and according to the example described above, the input block comprises only the ECH slicing module, and no registration or residual transformation is performed on the candidate and reference images prior to their slicing. However, it is possible to perform registration and / or residual transformation on the reference and candidate images prior to their slicing.

[0126]

[0101] According to a second variant of the second embodiment of the invention, the same neural network CNN' is implemented as before, but its learning, training, is carried out with differently constituted training sets.

[0127]

[0102] Thus, according to the second variant of this second embodiment, the neural network is trained with two balanced training sets: a first set of image tags labeled "Glass-direct" and a second set of image tags labeled "non-Glass". This labeling may have been performed by human operators and / or automatically by implementing, for example, the first embodiment of the invention dealing with images resulting from the superposition of texture images with a random component.

[0128]

[0103] Each "Glass-direct" image in the first set consists of the concatenation of two image tiles comprising at least one texture with a random component, the two tiles exhibiting, after superposition, the Glass phenomenon for an observer with normal visual acuity without the need for any prior processing or operation, or even during the superposition. This will be referred to as simple superposition.

[0129]

[0104] Each "non-Glass" image in the second set is made up of the concatenation of two image tiles comprising at least one texture with a random component, the superposition of the two tiles not being able to cause the Glass phenomenon even after a fine registration operation followed by the application of a residual transformation.

[0130]

[0105] Once the learning has been carried out with the two sets "Glass direct" and "non-Glass", the CNN' neural network obtains in the exploitation phase a success rate of more than 90%.

[0131]

[0106] Training according to the second variant makes it possible to obtain a CNN' neural network less sensitive to image acquisition and registration conditions, which makes it more robust and facilitates its implementation in the context of industrial and / or consumer applications for the authentication of mass-produced material subjects.

[0132]

[0107] In the preceding example, two categories or classes of imagelets were identified: "Glass-direct" imagelets and "non-Glass" imagelets. It is possible to identify a third category of imagelets called "Glass-indirect." Each "Glass-indirect" imagelet consists of the concatenation of two image tiles, each containing at least one texture with a random component. The simple superposition of these two tiles does not exhibit the Glass phenomenon to an observer with normal visual acuity. In contrast, the superposition of the two tiles of the "Glass-indirect" imagelet exhibits a Glass pattern, either after coarse registration or after fine registration followed by a residual transformation. The imagelets of the specific categories "Glass-direct" and "Glass-indirect" both belong to the general category of imagelets.

[0133] "likely to cause the appearance of Glass patterns."

[0134]

[0108] In the context of this application, it may be said that the blocks of a small image

[0135] "Direct glass" are normalized or normalized with respect to each other. This normalization can be carried out by an operator or semi-automatically with automatic registration and control by an operator of the superposition, or completely automatically with automatic registration and the automatic application of a predefined residual transformation.

[0136]

[0109] Surprisingly, the CNN' neural network trained according to the second variant of the second embodiment proves capable of identifying certain "glass-indirect" pairs or imagelets as "recognized," meaning that it is able to identify non-normalized imagelets or pairs of imagelets even though the transformation from one to the other is outside the domain of residual transformation. This characteristic is particularly advantageous because it allows, in practice, the implementation of a coarse registration, which is faster and less resource-intensive than a fine registration.Similarly, the CNN' neural network trained according to the example of the second variant is able to classify in the recognized category a pair of identical images which belong to the category "Glass-indirect" insofar as their simple superposition cannot make a Glass pattern appear but it is possible to make the Glass pattern appear by superimposing the two images one of which will have undergone a residual transformation.

[0137]

[0110] During the operation of the trained neural network CNN', the input block preferably includes, upstream of the slicing module ECH, a registration module 40 that ensures the registration of the reference image Ref and the candidate image Cand relative to each other. The INPUT block then includes a transformation module 41 that induces a residual transformation relative to each other on the images supplied to it. To this end, the transformation module 41 applies, for example, the residual transformation to one of the two images and not to the other. However, the transformation module 41 could operate differently insofar as the final result of the processing it performs corresponds to a residual transformation relative to each other of the two images supplied to it.As an example, the residual transformation is a rotation of a few degrees combined with a rectilinear translation of about ten pixels. According to the example illustrated in Figure 5, the transformation module 41 is located between the registration module 40 and the clipping module ECH.

[0138] However, it could be considered to place the transformation module 41 after the ECH slicing module. The transformation module 41 then operates on the reference tiles Ref and candidate Cand before their concatenation, which is performed by a concatenation module not shown.

[0139]

[0111] In the context of the invention and of the present application, the terms images, imagettes and tiles refer to objects, in the mathematical and computer science sense of the term, of the same nature, so that a treatment described in the context of the present application as being applied to one of these three types of object can, mutatis mutandis, be applied to the other two objects depending on the time when said treatment is implemented in the course of the process according to the invention.

[0140]

[0112] In the context of the examples described above, the labels used to distinguish the training sets are "recognized", "not recognized", "Glass-direct", "Glass-indirect" and "Non-Glass" to facilitate the disclosure of the invention, but other labels could be used insofar as they allow the types of training sets to be distinguished from each other.

[0141]

[0113] Preferably, the image sets used in the training sets include a large number of image sets with random textures from different types of material. This configuration of the training sets allows for a more efficient CNN neural network, as it is capable of identifying Glass patterns on a large number of superimposed image sets with random textures from a wide variety of subjects. Furthermore, the CNN thus trained will be able to recognize a Glass pattern resulting from the superposition of random textures that were not present in the training sets.

[0142]

[0114] It should be noted that an artificial neural network exhibits, during operation, a behavior that reflects the learning stages it has undergone. Thus, an artificial neural network configured to be sensitive to Glass patterns in the context of images with random texture components will be recognizable among others, particularly by subjecting it to a set of tests containing such images and analyzing the responses provided by the artificial neural network under study.

[0143]

[0115] Within the framework of the invention, the labeling of training images or image sets can be carried out in different ways. Initially, the labeling of the training sets is performed by an operator. This is referred to as supervised training. However, after an initial supervised training phase, it is possible to perform a second, unsupervised automatic training phase in which the image labeling is carried out by a CNN neural network that has undergone initial supervised training.

[0144]

[0116] Similarly, the image sets belonging to the training sets can be derived from textured images with a random component of material subjects, called natural image sets, or be image sets resulting from the superposition of synthetic images, called synthetic image sets. According to the invention, the training sets can comprise only natural image sets, only synthetic image sets, or a mixture of natural and synthetic image sets. Synthetic images are very well suited to simulating different lighting conditions (orientation, types, etc.), which can be very powerful for preparing a CNN for different shots of the candidate image and increasing its robustness.

[0117] In this regard, in order to increase the performance and robustness of the method according to the invention, some or even all of the "Glass" or "Glass-direct" type training pairs,

[0145] "Glass-indirect" images consist of two images of the same texture with a random component. These images differ from each other in that they are not captured under the same conditions. Among the parameters defining these conditions are: the image acquisition sensor, the lens associated with said sensor and its settings, the shooting angle, and the lighting conditions during capture, although this list is not exhaustive. In the context of this invention, the shooting conditions differ in that at least one of these parameters is not identical for the two images, while of course not preventing the appearance of the Glass pattern when the two images are superimposed after registration and / or residual transformation, if applicable.

[0146]

[0118] In the context of an example of the implementation of synthetic imagery for training, the applicant advantageously used randomly generated Perlin texture images to constitute the training sets of a CNN' neural network according to the invention which, after training, proved to be effective in the authentication of paper-type material subjects.

[0147]

[0119] Furthermore, training sets can be created by one or more operators through manual or supervised semi-automatic operations. However, the creation of training sets could be fully automated. Thus, training sets of "Glass-direct" imagelets or imagelet pairs can be automatically generated from series of images with a continuous random texture component of various material subjects. The generation of each "Glass-direct" imagelet pair consists of extracting one imagelet from an image in said series, which forms the first image of the pair, and the second imagelet of the pair is formed by applying a residual transformation to the first.For "Glass-indirect" type pairs, the process can be carried out in the same way as for "Glass-direct" type pairs, with the difference that instead of applying a residual transformation to the first image to constitute the second image, a rotation of between 90° and 270° is applied to the first image to form the second image.

[0148]

[0120] Similarly, training sets of "non-Glass" imagelets or imagelet pairs can be automatically generated from series of continuously random textured images of various material subjects. The generation of each "non-Glass" imagelet pair consists of extracting an imagelet from a textured image of two distinct subjects, or an image from two distinct areas of the same subject.

[0149]

[0121] As already mentioned, the automation of training and / or the constitution of training sets can also result from the implementation of the first embodiment of the invention or even the second embodiment.

[0150]

[0122] Furthermore, within the scope of this application and the context of the invention, the terms blocks and modules may refer to purely software elements, purely hardware elements, or elements combining software and hardware implementations. This also applies to the implementation of the entire invention, which may be purely software-based, purely hardware-based, or a combination of software and dedicated hardware components.

[0151]

[0123] It should also be indicated that within the framework of the invention the input block can implement various types of calculator processes to carry out the different processing it performs, that it can in particular implement specific artificial neural networks configured to perform the operations of registration, cutting and application of transformations to the images to be processed.

[0152]

[0124] Furthermore, in the examples described above, the images to be processed are segmented before being provided to each neural network configured to be sensitive to Glass patterns. However, such a mode of operation is not strictly necessary for the realization of the invention, since it is possible to implement artificial neural networks sized for processing images from acquisitions without it being necessary to segment these images beforehand.

Claims

Demands 1. A method for the unitary authentication of a material subject consisting of confronting, on the one hand, at least one so-called reference image of at least one authentication region of an authentic subject, the reference image comprising at least one texture with a random component, and on the other hand, a so-called candidate image of at least one authentication region of a candidate subject, the candidate image comprising at least one texture with a random component, characterized in that it implements at least one artificial neural network and in that the images are processed jointly by the same artificial neural network which is configured to be sensitive to Glass patterns in the context of images with a texture with a random component.

2. Authentication method according to claim 1, characterized in that the artificial neural network has been trained with at least one training set comprising training pairs, at least some of which comprise two images whose superposition reveals or is likely to reveal a pattern analogous to a Glass pattern.

3. Authentication method according to claim 1 or 2, characterized in that for their joint processing the candidate and reference images are superimposed and the artificial neural network is provided with an image resulting from the superposition of the candidate image and the reference image, the artificial neural network having been trained using images with random texture components, some of which have a Glass pattern.

4. Authentication method according to claim 1 or 2, characterized in that for their joint processing the two images to be processed, either from each training pair, or the reference and candidate images, are assembled by being juxtaposed without overlap to form only one and the same image which will be provided to the artificial neural network.

5. Authentication method according to claim 1 or 2, characterized in that for their joint processing the two images to be processed are provided simultaneously to the same artificial neural network and / or to two distinct inputs of the same artificial neural network.

6. Authentication method according to any one of the preceding claims, characterized in that the artificial neural network is adapted to process images of given dimensions and when the images to be processed are larger, the images to be processed are cut into sub-images or imagelets of suitable size and constituted in pairs of sub-images or imagelets to be processed jointly which are submitted to the artificial neural network.

7. Authentication method according to any one of the preceding claims, characterized in that prior to their joint processing the images to be processed are recalibrated with respect to each other.

8. Authentication method according to claim 7, characterized in that, subsequent to registration and prior to their joint processing, the images to be processed undergo a residual transformation relative to each other.

9. Authentication method according to any one of the preceding claims, characterized in that, after registration, the images to be processed are cut into sub-images or imagelets of suitable size and the sub-images or imagelets constituting each pair to be processed are subjected to a residual transformation with respect to each other, prior to their joint processing.

10. Authentication method according to claim 3 or 4 or 5, characterized in that, within the framework of the network training phase, the images of each training pair are recalibrated against each other prior to their joint submission to the artificial neural network.

11. Authentication method according to claim 10, characterized in that after registration and before supplying to the artificial neural network the images of each training pair are subjected to a residual transformation with respect to each other.

12. Authentication method according to any one of the preceding claims characterized in that each candidate image is taken from a video stream.

13. Authentication method according to any one of the preceding claims, characterized in that the artificial neural network is a convolutional neural network.

14. Authentication method according to claim 4 characterized in that it comprises a learning phase of the artificial neural network by means of training pairs labeled by the method according to claim 3.

15. Authentication method according to any one of the preceding claims, characterized in that after implementation of the artificial neural network, it comprises a step of presenting to a user the superimposed reference and candidate images.

16. Product computer program comprising code instructions for executing a method according to any of the preceding claims of authenticating a material subject from a reference image and a candidate image, when the program is executed on a computer.

17. A computer-readable storage means on which a computer program includes code instructions for executing a method according to any one of claims 1 to 13 for authenticating a material subject from a reference image and a candidate image.

18. Computer device comprising at least display means, image acquisition means, user input means, information storage means communicating with computing and control means configured to implement the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Method for authenticating and / or checking the integrity of a subject

    EP3380987A1

  • Method of augmented authentification of a material subject

    US10990845B2

  • Non-counterfeitable document system

    US4423415A

  • procedure D'AUTHENTIFICATION PAR MOTIF DE GLASS

    FR3044451A3