Distillation neural network system, improved training methods
The student-teacher neural network system addresses the inefficiencies of existing defect detection methods by minimizing cost functions based on linear processing layers, resulting in faster and more flexible learning suitable for industrial applications.
Patent Information
- Application Number
- FR2023002204
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-03-09
Smart Images

Figure 00000018_0000 
Figure 00000018_0001 
Figure 00000019_0000
Abstract
Description
Title of the invention: Distillation neural network system, improved training methods Technical field
[0001] The invention lies in the field of artificial neural networks. More particularly, the invention relates to the use of such neural networks to detect anomalies or defects in manufactured products, for example at the time of their manufacture, or in raw products. Prior art
[0002] The automatic detection of defects or anomalies on manufactured products, from images or videos of these products provided as input to an electronic system, is of the greatest interest, particularly in the context of quality control processes implemented in industrial environments.
[0003] In this field, computer vision and artificial intelligence techniques are increasingly being used to provide automatic fault detection and localization solutions that are as reliable as possible and relatively inexpensive. Many of these solutions rely on the use of artificial neural networks, which, once trained, prove to be particularly efficient and well-suited to carrying out specific applications of this type, involving pattern recognition and / or data classification.
[0004] [Fig.l] schematically and simplifiedly presents a standard prior art neural network architecture, in which a variable number of layers of neurons follow one another to produce an output S (for example a classification) as a function of input data E injected into the network (for example images of a machined part or a product). Such a network generally comprises an input layer CE and an output layer CS, between which one or more intermediate layers (Cil, CI2, CI3) are arranged.
[0005] The intermediate layers each comprise one or more neurons (not shown), the neurons of a layer taking their inputs from the outputs of the neurons of the previous layer. Within a neuron, the inputs of the neuron are combined linearly (i.e. via a weighted sum of the inputs), and the result of this linear combination of the inputs is injected as input to a non-linear activation function (e.g. a sigmoid or ReLu function) which delivers the output value of the neuron which will in turn feed the next layer of the network. In a simplified manner, and as illustrated in [Fig.l], we can therefore consider that each intermediate layer (Cil, CI2, CI3, etc.) can be decomposed into a processing layer linear L (linear combination), followed by a nonlinear processing layer NL (nonlinear activation).
[0006] The different weights of the neural network (i.e. the different weighting coefficients used for implementing the linear combination of the inputs at each neuron) are adjusted during a learning phase, by minimizing a cost function. More particularly, during this learning phase, an output generated by the network is compared with an expected output provided by the user (supervised learning) or generated automatically (unsupervised learning), and the weights of the neural network are adjusted progressively, using a gradient descent type backpropagation algorithm, so as to reduce the error (i.e. minimize the cost function) between the output obtained and the expected output.By repeating this procedure a sufficient number of times with many training examples, the network's weights are gradually adapted until the generated output matches the ground truth as closely as possible in most cases. The neural network can then be used to perform the specific application for which it was thus trained, such as pattern recognition or data classification for example.
[0007] Such deep learning-based neural networks for image analysis have shown great potential in recent years in these fields relating to image classification and object detection. Based on a large amount of images marked or labeled with object classes or concepts, artificial intelligence algorithms based on neural networks are notably capable of learning image representations with great generalization power, and of providing detection performances close to those of humans.
[0008] Although promising, particularly in the field of defect detection, numerous obstacles mean that existing solutions based on standard neural networks (i.e. with an architecture of the type illustrated in [Fig.l]) remain poorly adapted to the real needs and constraints of industrial players. In particular, when trying to deploy such solutions in practice in a production environment, the large amount of data required to train the network, the need to identify and label samples en masse for each type of defect, the slowness of learning, and the rigidity of current technology in the face of constantly changing situations are among the main bottlenecks that prevent these solutions from reaching their full maturity for applications on an industrial scale.
[0009] Various solutions have been explored to at least partially address some of these problems. For example, instead of learning to detect faults, a different, more powerful approach is to train the neural network to learn the normal (i.e., defect-free) appearance of objects, so as to then detect any deviation from normality such as a defect or anomaly. Through this paradigm shift, the amount of data required to train the network is significantly reduced and this data is easier to gather (since this time we will target only "good" products, i.e., "normal" images of parts, without defects), and learning is faster (since the system does not need to learn to detect several classes of defects). However, existing solutions based on this approach for anomaly detection still do not allow certain important constraints identified previously to be lifted.In particular, they still rely on complex deep learning schemes that remain inflexible and require substantial learning effort, making them still poorly suited for implementation in real industrial production conditions.
[0010] In an attempt to overcome these drawbacks, solutions based on the concept of knowledge distillation and on student-teacher type neural network systems have been developed and used in the field of fault detection, with the particular aim of reducing the complexity of the neural network models used (network compression technique) and thus making them a little more suited to the constraints of manufacturers. [Fig.2] presents in a schematic and simplified manner an example of the architecture of such a student-teacher type neural network system, according to the prior art, which is based on the use of a pair of neural networks: a REN teacher network and a REL student network.In such an approach, the REN teacher network is pre-trained: in other words, it has already been trained with a large amount of data (for example, based on a dataset of the ImageNet type, well known in the field of computer vision, comprising millions of arbitrary images annotated in order to specify the presence or absence of certain classes of objects in each image) to carry out a task, for example an image classification task, which ultimately has little importance (such a task is therefore sometimes referred to as a "proxy" task). In a system of student-teacher networks, the idea is not to exploit the teacher network for its ability to carry out a particular task, but rather to exploit its capacity to extract at the level of its intermediate layers (and more particularly the intermediate layers Cil, CI2, etc., upstream of the network) of the important generic characteristics of the images provided as input to this network. The final decision layers of the REN teaching network are therefore not significant in this framework.
[0011] By integrating a pre-trained REN teaching network into a neural network system, it becomes possible, during a learning phase, to force the REL student network to adapt to given characteristics, typically characteristics representative of the normality of a product, even if the student network has a significantly smaller number of intermediate layers (Cil', CI2', etc.) than the teacher network. To do this, during the learning phase, IMG images of defect-free products are injected into the system's input, and the features extracted at the output of the intermediate layers of the REN teacher network, which are then generic characteristics representative of a "normal" (i.e. defect-free) product, are used to train the REL student network, by minimizing FC cost functions between the corresponding intermediate layers (Cil and Cil', CI2 and CI2', etc.) of these two networks, via a progressive adjustment of the weights of the REL student network.In this way, the student network also learns not to solve a specific task, but to specialize in the "normal" (i.e., defect-free) appearance of an object.
[0012] Subsequently, in the operational phase (i.e. once the learning phase is completed), the injection into the neural network system of an image of a product presenting defects or anomalies (i.e. of a product which is not “normal”) will tend to produce different characteristics in the two networks at the level of their corresponding intermediate layers, a measurement of these differences thus allowing the detection of the presence or on the contrary the absence of the defects or anomalies in question.
[0013] The previously described approach, based on the use of a prior art student-teacher type neural network system as illustrated in relation to [Fig.2], still suffers from certain drawbacks which may prove prohibitive for use on an industrial scale. Firstly, although faster than traditional deep learning approaches, the learning is still not fast enough for certain industrial applications, in particular because it remains based on standard backpropagation and gradient descent algorithms, i.e. non-linear optimization implemented via highly iterative and therefore relatively slow processes. Secondly, the learning of the student network is always carried out before the deployment of the system in a production environment (i.e.offline), and the system actually deployed therefore has a fixed dimension, which limits the system's ability to adapt to continuous changes that are nevertheless frequent in an industrial environment. Thirdly, a supervision step is always necessary, during which normal samples are presented to the system, before the system can be deployed in production and effectively used for fault detection.
[0014] There is therefore a need for neural network-based solutions that are more efficient, faster and more flexible than those of the prior art, particularly in the field of fault detection. Summary of the invention
[0015] The present technique makes it possible to propose a solution aimed at remedying certain drawbacks of the prior art. According to one aspect, the present technique relates in fact to a system of neural networks of the student-teacher type, in which a first pre-trained neural network, called the teacher network, is used during a learning phase to train a second neural network, called the student network, by minimizing at least one cost function between corresponding intermediate layers of said networks, said intermediate layers each comprising a linear processing layer followed by a non-linear processing layer. According to the general principle of the proposed technique, said system is configured so that, during said learning phase, said cost functions are minimized as a function of data obtained directly at the output of the linear processing layers of said corresponding intermediate layers.
[0016] In a particular embodiment, said neural network system is further configured so that, during said learning phase, the linear processing layers of the intermediate layers of the student network have as inputs the outputs of the immediately preceding intermediate layers of said teacher network.
[0017] In a particular embodiment, said neural network system is used to detect the presence or absence of a defect on a target product, from at least one image of said target product provided as input to said system.
[0018] According to a particular characteristic of this embodiment, during said learning phase, only images of products of the same type as the target product in which said products do not present defects are injected into the input of said system.
[0019] Alternatively, according to another particular characteristic of this embodiment, during said learning phase, any images of products of the same type as the target product are injected into the input of said system, and said system integrates robust estimation means configured to automatically reduce, during said learning phase, the sensitivity of said system to images of products comprising a defect among said any images.
[0020] In a particular embodiment, the dimensioning of the student network is adapted incrementally, by successively adding, one after the other, intermediate layers to said student network, as long as a predetermined efficiency threshold has not been reached in the detection of the presence or absence of defects on target products.
[0021] According to another aspect, the proposed technique also relates to a method of adapting the sizing of a student network of a neural network system as described previously in its various embodiments. Such a method comprises the following steps:
[0022] (a) adding an intermediate layer to said student network;
[0023] (b) estimation of the parameters associated with said added intermediate layer, of independently of the rest of the said student network;
[0024] (c) obtaining a measurement representative of a precision achieved by the system of neural networks, said measure taking into account all the intermediate layers added to said student network;
[0025] (d) comparison of said measurement representative of an accuracy achieved with a threshold of predetermined precision, and
[0026] - when said measurement representative of an achieved precision is greater than said predetermined precision threshold, stopping said student network sizing adaptation process;
[0027] - when said measurement representative of an achieved precision is less than said predetermined accuracy threshold, implementing a new iteration of steps (a), (b), (c) and (d).
[0028] According to another aspect, the proposed technique also relates to a computer program product downloadable from a communication network and / or stored on a computer-readable medium and / or executable by a microprocessor, comprising program code instructions for executing a method for adapting the sizing of a student network of a neural network system as described above, when executed on a computer.
[0029] The proposed technique also relates to a computer-readable recording medium on which is recorded a computer program comprising program code instructions for executing the steps of the method as described above, in any of its embodiments.
[0030] Such a recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a USB key or a hard disk.
[0031] On the other hand, such a recording medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means, so that the computer program contained therein is remotely executable. The program according to the invention may in particular be downloaded over a network, for example the Internet.
[0032] The different embodiments mentioned above can be combined with each other. for the implementation of the invention. Figures
[0033] Other characteristics and advantages of the invention will appear more clearly on reading the following description of a preferred embodiment, given as a simple illustrative and non-limiting example, and the appended drawings, among which:
[0034] [Fig.l] presents in a schematic and simplified manner a standard neural network architecture according to the prior art;
[0035] [Fig.2] presents in a schematic and simplified manner an architecture of a teacher-student type neural network system, according to the prior art;
[0036] [Fig.3] presents in a schematic and simplified manner a first improved configuration of a teacher-student type neural network system, in the learning phase, in a particular embodiment of the proposed technique;
[0037] [Fig.4] presents in a schematic and simplified manner a second improved configuration of a teacher-student type neural network system, in the learning phase, in a particular embodiment of the proposed technique;
[0038] [Fig.5] illustrates the different stages of a method for adapting the dimensioning of a student network of a neural network system, in a particular embodiment of the proposed technique;
[0039] [Fig.6] presents a simplified architecture of a device for adapting the sizing of a student network of a neural network system, in a particular embodiment. Detailed description of the invention
[0040] The present technique makes it possible to overcome some of the aforementioned drawbacks.
[0041] The proposed technique aims in particular to propose a neural network system which is at the same time efficient, fast and flexible, and therefore well suited to being deployed in industrial production environments, for example for applications of automatic detection of defects on raw or manufactured parts or products.
[0042] According to one aspect, the present technique relates more particularly to innovative configurations of student-teacher type neural network systems, which notably allow the simplification and optimization of the learning of the student network.
[0043] In all the figures of this document, elements and steps of the same nature are designated by the same reference.
[0044] As already described in connection with [Fig. 2] of the prior art, and again presented below in connection with Figures 3 and 4 illustrating various particular embodiments of the present technique, a pupil-type neural network system A teacher is a system in which a first pre-trained neural network, called a teacher network REN, is used during a learning phase to train a second neural network, called a student network REL. Independently of the intended application for the neural network system, the teacher network has for example been previously pre-trained for any classification task, using a data set comprising a very large number (e.g. several tens or hundreds of thousands, or even millions) of annotated images of various entities or classes of objects (objects, living beings, etc.).The implementation of such pre-training of the teaching network is not the subject of the present technique, but it results therefrom, as already described in relation to the prior art, that the teaching network thus already has the capacity to extract at the level of its intermediate layers (and more particularly of the intermediate layers upstream of the teaching network) important generic characteristics of the images provided as input to this network.
[0045] The neural network system according to the present technique is however distinguished from the systems of the prior art in that it is configured in a particular manner, in particular during the learning phase, according to innovative configurations giving it multiple advantages.
[0046] According to a first configuration, illustrated in relation to [Fig. 3] in a particular embodiment, rather than minimizing during the learning phase the FC cost functions as a function of the data obtained at the output of the non-linear processing layers of the intermediate layers of the network, as is conventionally carried out in the systems of the prior art such as presented in [Fig. 2], it is proposed to carry out this operation as a function of data obtained upstream, directly at the output of the linear processing layers of the intermediate layers. In other words, the adjustment of the weights of the intermediate layers of the student network is carried out so as to adapt to the output of the linear processing layers of the intermediate layers of the teacher network, and no longer of the non-linear processing layers as in the existing system.
[0047] Such an adaptation is possible on the one hand due to the use of two networks (the teacher network and the student network) operating in parallel, and on the other hand due to the fact that the activation functions conventionally used for the implementation of the non-linear processing layers of the intermediate layers - in particular in systems used for fault detection applications - are increasing continuous monotonic functions.
[0048] Thus, minimizing a cost function as a function of output data from the linear processing layer of the intermediate layers is, from the overall point of view of the learning purpose of the student network, equivalent to minimizing a cost function based on the output data of the nonlinear processing layer of the intermediate layers as performed in the prior art.
[0049] On the other hand, from the point of view of the process of optimizing the weights of the REL student network during its learning phase, such a configuration according to [Fig. 3] is particularly advantageous in that it makes it possible to switch from an objective of solving a non-linear optimization problem as posed with the solutions of the prior art to an objective of solving a purely linear optimization problem within the framework of the proposed technique. In other words, a previously complex optimization problem (non-linear minimization) is transformed, by means of the present technique, into a much less complex optimization problem of determining a linear regression model with respect to the parameters of the linear layers of the system.Simple and fast linear estimation techniques, for which a solution is otherwise directly accessible without the need to resort to a process of successive iterations as is the case for solving a non-linear problem, can then be implemented. For example, the present technique makes it possible to use tools such as Singular Value Decomposition (or SVD) for this purpose instead of the classic gradient descent backpropagation algorithms conventionally implemented in existing networks. Such linear estimation tools provide exact solutions quickly and also make it possible to guard against certain common problems of existing neural networks, such as the problem of vanishing gradients for example.
[0050] A second configuration, presenting additional and complementary adaptations to the first configuration previously presented, is illustrated in relation to [Fig.4], in a particular embodiment of the proposed technique.
[0051] In this second configuration, it is proposed to exploit in an even more in-depth manner the parallelism between the teacher network REN and the student network REL, by directly injecting into the input of the intermediate layers of the student network the output data of the immediately preceding intermediate layers of the teacher network, during the learning phase of the student network. In other words, as shown in [Fig.4], the neural network system is configured in such a way that an intermediate layer of rank n of the student network REL accepts as input no longer the output of the intermediate layer of rank n-1 of the student network REL, but that of rank n-1 of the teacher network REN (for example, the output of the non-linear processing layer NL of the intermediate layer Cil of the teacher network REN is injected into the input of the linear processing layer L of the intermediate layer CI2' of the student network, the output of the non-linear processing layer NL of the intermediate layer CI2 of the teacher network REN is injected into the input of the linear processing layer. L of the intermediate layer CI3' of the student network, etc.).
[0052] Such an adaptation is based on the fact that the propagation of data from the input to the output of a neural network corresponds to the implementation of a composition of functions, and that minimizing the cost function at the output of each linear processing layer separately is equivalent, due to the continuity properties of the functions used, to minimizing the composition of the functions (the limit of a composition of functions is the composition of the limits, given the continuity).
[0053] With this configuration in the learning phase, a complete decoupling of the estimation of the weights associated with the different linear processing layers of the intermediate layers (Cil', CI2', etc.) of the student network is obtained, insofar as
[0054] - the input of a linear processing layer L of an intermediate layer of the network student REL is taken directly from the output of a previous intermediate layer of the teacher network REN (except for the very first intermediate layer of the student network);
[0055] - the minimization of the cost function is carried out at the output of the layer of linear processing of the intermediate layer of the student network, based on output data from the linear processing layer of the corresponding intermediate layer of the teacher network (according to the general principle already described in relation to the first configuration presented in relation to [Fig.3]).
[0056] In other words, in the configuration according to [Fig.4], the optimization of each intermediate layer (Cil', CI2', CI3', etc.) of the student network becomes a problem independent of the rest of the student network, which does not depend in particular on the optimization of the preceding or following intermediate layers of the student network.
[0057] Such a configuration is particularly interesting in many respects.
[0058] Firstly, it is no longer necessary to use error backpropagation algorithms to estimate the weights of each intermediate layer of the student network.
[0059] Secondly, each intermediate layer of the student network can be optimized independently of the other intermediate layers of this network, which makes it possible to parallelize the optimization operations of the different layers of the student network during the learning phase, with the effect of an overall speed gain.
[0060] Third, learning can be easily distributed for each intermediate layer of the student network since these intermediate layers are no longer interdependent.
[0061] Fourthly, the direct use of the characteristics of the teacher network as input to the linear processing layers of the intermediate layers of the student network makes it possible to accelerate the convergence towards the solution making it possible to minimize the function cost, when an iterative optimization method is used. In particular, convergence is much faster than in standard architectures such as the one illustrated in [Fig.2], in which many iterations are required in the training phase before obtaining relevant inputs on all the intermediate layers of the student network. Indeed, in existing systems, the input data injected into an intermediate layer of the student network correspond to the output data of the previous layer of this network. Also, as long as the parameters of a layer are still poorly estimated, the results produced at the output of this layer are not relevant and are therefore similar to noise injected into the following layer.In other words, as long as the process of adjusting the weights of a layer of the student network has not reached a certain degree of adaptation to the characteristics of the corresponding layer of the teacher network, the adjustment of the weights of the next layer of the student network is not carried out on relevant data. The overall process of adjusting all the weights of the student network is therefore slow to converge, in existing solutions. The configuration according to the present technique, as illustrated in relation to [Fig.4], does not have such drawbacks: in fact, as the teacher network is already pre-trained, the output data of the linear processing layers of the intermediate layers of the teacher network which are injected as input to the intermediate layers of the student network are already relevant data (the process of adjusting the weights of the teacher network having already been completed during the prior training of this network).The learning phase of the student network is therefore faster with the proposed technique than in existing solutions.
[0062] A neural network system according to the proposed technique is for example particularly suitable for use in applications for detecting the presence (or absence) of defects or anomalies on target products, from at least one image of the target product provided as input to said system. Such images can for example be provided by image acquisition devices (e.g. photographic device, camera) installed on a production line (acquisition of images of products conveyed on a conveyor belt, for example).
[0063] In a particular embodiment, during the learning phase of the student network of the neural network system, only images of products of the same type as the target product, which do not have defects, are presented as input to the system. Thus, in this “supervised” type learning mode, the student network is allowed to specialize in learning the “normal” appearance of the product considered (i.e. without defects), to the extent that the generic characteristics extracted at the level of the intermediate layers of the pre-trained teaching network will then be representative of such normality of a product. The properties of layer independence and rapid convergence of the optimization presented previously are then particularly interesting, in that they make it possible to limit the number of iterations during which the characteristics extracted at the level of the intermediate layers and used for the optimization are not yet relevant, that is to say to limit the optimization on the basis of data which would not yet be representative of the "normality" of the product considered.
[0064] In a particular alternative embodiment, it is also possible to implement “unsupervised” type learning, by integrating robust estimation means into the neural network system. In this case, during the learning phase of the student network of the neural network system, any images of products of the same type as the target product are injected into the system, with the assumption that a majority of the images thus injected are images of “normal” products, i.e. not exhibiting defects or anomalies. In other words, it is considered that statistically, on a panel of products of the same type selected at random, the vast majority of products (e.g. more than 95% of the products) will have a normal appearance, while only a small minority of products (e.g. less than 5% of the products) will exhibit a defect.The robust estimation means are then configured to reduce the sensitivity of the system to product images comprising a defect among the images injected as inputs, based for example on statistical anomaly detection techniques. Product images not corresponding to normality can thus for example be automatically rejected or discarded during the learning phase. Such a variant is of definite interest, in that it allows learning to be implemented even after deployment of the system in a production environment, and moreover in a fully unsupervised mode.In other words, it is sufficient, for example, for products to "scroll" for a certain time (considered as a learning time) under the image acquisition device coupled with the neural network system, under normal production conditions, for the learning phase and therefore the optimization of the system to be carried out automatically and autonomously.
[0065] The decoupling during the learning phase of the different intermediate layers of the student network as obtained with the neural network system configuration presented in [Fig.4] offers further particularly interesting advantages. Thus, it makes it possible in particular to easily adapt the dimensioning of the student network to the different needs of manufacturers and / or to changes in the industrial environment in which the solution is deployed. More particularly, this decoupling facilitates the progressive addition of complexity to the model, by making possible an “à la carte” or “on demand” adaptation of the number of intermediate layers of the student network, even when the system is already deployed in an environment of production. Indeed, since the parameters of the intermediate layers of the student network are estimated independently for each layer (i.e. without any interdependence between the different intermediate layers of the student network), layers can be added (or removed) as needed, without this modification affecting the other layers of the student network. A criterion such as reaching a predetermined efficiency threshold (for example, an efficiency threshold in detecting the presence or absence of defects on target products), or any other means of decision, can, for example, be used to determine whether or not it is useful to add an intermediate layer to the student network.
[0066] [Fig.5] shows, in a particular embodiment of the proposed technique, an example of a method for adapting the sizing of a student network of a neural network system configured in the embodiment described previously in relation to [Fig.4]. Such an example is given purely for illustrative and non-limiting purposes, in the context of an application to the detection of defects on a target product. Such a method is for example implemented after deployment of the neural network system in a production environment, possibly with a student network in which no intermediate layer is yet instantiated, and it comprises the following steps:
[0067] (a) a step 51 of adding an intermediate layer to the student network (at the end of the very first implementation of this step, the student network includes for example only one intermediate layer);
[0068] (b) a step 52 of estimating the parameters (i.e. the weights) associated with the layer in added intermediate, using the approach already described in relation to [Fig.4] (in other words, thanks to the particular configuration of the neural network system during the learning phase, this estimation is carried out independently of the rest of the student network, and in particular of the other intermediate layers possibly already present);
[0069] (c) a step 53 of obtaining a MES measurement representative of a precision achieved by the neural network system, said measurement taking into account all the intermediate layers added to said student network (such a measurement is for example determined based on the output characteristics of the intermediate layers of the student network, with regard to the characteristics of the corresponding intermediate layers of the teacher network, for example by evaluating the values of the different cost functions between corresponding layers);
[0070] (d) a step 54 of comparing the MES measurement representative of a precision reached with a predetermined precision threshold SL and:
[0071] - when said measurement representative of an achieved precision is greater than the threshold of predetermined precision, stopping said process of adapting the dimensioning of the student network (in other words, the number of intermediate layers of the student network is considered sufficient for the desired objective and there is no need to add any more, the sizing adaptation process is then complete);
[0072] - when said measurement representative of an achieved precision is less than said predetermined accuracy threshold, implementing a new iteration of steps (a), (b), (c) and (d) (in other words, the level of accuracy achieved is still considered insufficient, and it is necessary to add at least one more intermediate layer to the student network to improve its accuracy).
[0073] The neural network system according to the present technique is therefore particularly flexible.
[0074] Once the learning phase is complete, the neural network system according to the proposed technique can be used in an operational phase. For example, for a defect detection application as described throughout this document, injecting into the neural network system an image of a product exhibiting defects or anomalies (i.e., a product that is not “normal”) will tend to produce different characteristics in the two networks at their corresponding intermediate layers, a measurement of these differences thus allowing the detection of the presence or, on the contrary, the absence of the defects or anomalies in question. It should be noted that for this operational phase (i.e., inference), the neural network system can be used in its configuration illustrated in relation to [Fig.4], or have been reconfigured after the learning phase to be used in a more classic configuration as illustrated in relation to [Fig.2].
[0075] According to another aspect, the proposed technique also relates to a device capable of carrying out the method of adapting the dimensioning of a student network of a neural network system presented previously. More particularly, such a device according to the present technique comprises means allowing the implementation of steps (a), (b), (c) and (d) previously described, ie:
[0076] - means for adding an intermediate layer to the student network;
[0077] - means for estimating the parameters associated with said intermediate layer added, independently of the rest of the said student network;
[0078] - means for obtaining a measurement representative of a precision achieved by the neural network system, said measurement taking into account all the intermediate layers added to said student network;
[0079] - means for comparing said measurement representative of an achieved precision with a predetermined accuracy threshold.
[0080] [Fig.6] represents, in a schematic and simplified manner, the structure of such a device, in a particular embodiment. The device according to the technique proposed comprises for example a memory 61 consisting of a buffer memory M, a processing unit 62, equipped for example with a microprocessor qP, and controlled by the computer program Pg 63, implementing steps of the method for adapting the dimensioning of a student network of a neural network system, according to at least one embodiment of the invention.
[0081] At initialization, the code instructions of the computer program 63 are loaded into the buffer memory before being executed by the processor of the processing unit 62.
[0082] The microprocessor of the processing unit 62 then carries out the steps of the method for adapting the dimensioning of the student network, according to the instructions of the computer program 63. More particularly, the processing unit 62 receives, for example, as input E a request for adapting the dimensioning of the student network of a neural network system, and proceeds to successively add, one after the other, intermediate layers to the student network, as long as the neural network system has not reached a predetermined precision threshold. Once this threshold has been reached, the processing unit 62 then delivers as output S a student-teacher type neural network system in which the number of intermediate layers of the student network is dimensioned for the execution, with the desired precision, of the application for which the student-teacher type neural network system is configured.
Claims
Claims
1. A student-teacher type neural network system, configured to detect the presence or absence of a defect on a target product from at least one image of said target product provided as input to said system, in which a first pre-trained neural network, called a teacher network (REN), is used during a learning phase to train, on the basis of images of products of the same type as the target product injected as input to said system, a second neural network, called a student network (REL), by minimizing at least one cost function (FC) between corresponding intermediate layers of said networks, said intermediate layers each comprising a linear processing layer (L) followed by a non-linear processing layer (NL), said system being characterized in that, during said learning phase,the linear processing layers of the intermediate layers of the student network have as inputs only the outputs of the immediately preceding intermediate layers of said teacher network, and in that, in response to the injection of an image of a product of the same type as the target product at the input of said system, said cost functions are minimized as a function of data obtained directly at the output of the linear processing layers (L) of said corresponding intermediate layers.,
2. System according to claim 1, characterized in that during said learning phase, only images of products of the same type as the target product in which said products do not present defects are injected into the input of said system.
3. System according to claim 1, characterized in that during said learning phase, any images of products of the same type as the target product are injected into the input of said system, and in that said system integrates robust estimation means configured to automatically reduce, during said learning phase, the sensitivity of said system to images of products comprising a defect among said any images.
4. System according to claim 1, characterized in that the dimensioning of the student network is adapted incrementally, by successive addition, one after the other, of intermediate layers to said student network, as long as a predetermined efficiency threshold has not been reached in the detection of the presence or absence of defects on target products.
5. Method for adapting the dimensioning of a student network of a neural network system according to claim 1, said method being characterized in that it comprises the following steps: (a) adding (51) an intermediate layer to said student network; (b) estimation (52) of the parameters associated with said added intermediate layer, independently of the rest of said student network; (c) obtaining (53) a measurement (MES) representative of a precision achieved by the neural network system, said measurement taking into account all of the intermediate layers added to said student network; (d) comparison (54) of said measurement (MES) representative of an accuracy achieved with a predetermined accuracy threshold (SL), and - when said measurement representative of an accuracy achieved is greater than said predetermined accuracy threshold, stopping said method of adapting the dimensioning of the student network; - when said measurement representative of an accuracy achieved is lower than said predetermined accuracy threshold, implementation of a new iteration of steps (a), (b), (c) and (d).
6. Computer program product downloadable from a communications network and / or stored on a computer-readable medium and / or executable by a microprocessor, characterized in that it comprises program code instructions for executing a method according to claim 5, when executed by a computer.