Method and device for processing multidimensional data for locating defects in a sample

The method transforms multidimensional data into a block of drops, calculates a structure tensor, and applies machine learning to accurately detect and localize defects in crystalline materials, addressing the challenge of precise defect localization in existing technologies.

FR3152186B1Active Publication Date: 2025-08-29COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2023008785
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-08-18
Publication Date
2025-08-29
Estimated Expiration
2043-08-18

AI Technical Summary

Technical Problem

Existing methods struggle to precisely locate defects such as dislocations and stacking faults in crystalline materials, which can cause deformations and malfunctions in devices, using conventional microscopy data analysis.

Method used

A method involving multidimensional data processing that includes transforming input data into a block of drops representing the atomic structure, calculating a structure tensor with directional gradients, forming a signature tensor, and applying machine learning classification to detect and localize defects like dislocations and stacking faults.

Benefits of technology

The method achieves precise localization of defects with improved accuracy, enabling early detection and prevention of material defects in devices, reducing computational resources and enhancing spatial precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000021_0000
    Figure 00000021_0000
  • Figure 00000022_0000
    Figure 00000022_0000
  • Figure 00000023_0000
    Figure 00000023_0000
Patent Text Reader

Abstract

Method and device for processing multidimensional data for locating defects in a sample This method for processing multidimensional data representative of a sample of crystalline material for locating defects comprises an acquisition of at least one image forming a block of input data, a transformation (34) of the block of input data to obtain a block of drops representative of an atomic structure of the sample, then, at at least one subset of points of said block of drops, calculation (36) of a structure tensor as a function of the values ​​of directional gradients, around each point P of the subset, calculation (38) of one or more characteristic values ​​of the structure tensor, and formation (50) of a signature tensor associated with the point P, the signature tensor having as components said calculated characteristic values,and then a detection and localization (54) of a defect by applying a machine learning classification method to the signature block bringing together the signature tensors associated with the points P of said drop block. Figure for the abstract: Figure 2,
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for processing multidimensional data for locating defects in a sample

[0001] The present invention relates to a method and a device for processing multidimensional data for locating defects in a sample of crystalline material, comprising an acquisition of at least one image forming a block of input data, the or each image of said block being representative of a part of the observed sample, said block of input data being represented in an N-dimensional space, N being greater than or equal to two.

[0002] The invention lies in the field of processing multidimensional data, obtained by observation of samples composed of one or more materials, for the analysis of their physical properties and their structures.

[0003] More particularly, the invention finds applications in the localization of defects observed in crystalline materials, for example applied in quality inspection in a production line of devices based on crystalline materials.

[0004] As is known, crystals are generally composed of an ordered mesh, defined by a Bravais lattice. By definition, they have a discrete diffractogram, i.e. composed of diffraction spots or balls, in the sense of the IUCR (International Union of Crystallography).

[0005] In practice, observed materials comprising one or more crystalline and / or amorphous species may exhibit defects, i.e. deformations or irregularities, which may induce stresses or malfunctions in devices using such materials, for example leakage currents in transistors.

[0006] For example, "dislocation" type deformations correspond to a linear defect of a crystal obtained by sliding of a part of the crystal along a plane. A dislocation is characterized by a slip vector, also called the Burgers vector of the dislocation. Known dislocations are "screw" dislocations, "corner" dislocations or mixed dislocations.

[0007] More generally, considering a data block in an N-dimensional space, the dislocation is defined by a sliding of a part of the data along a hyperplane. There are other types of defects inducing deformations that are sometimes very difficult to detect, such as the smallest stacking faults. The many other kinds of defects such as antiphase walls, twins, grain boundaries are also likely to cause local deformations of the materials.

[0008] It is useful to have high-performance tools for carrying out the inspection of materials, and detecting the possible presence of such defects, and in particular dislocations, in the calibration phase or in the technological optimization phase of various devices using the materials.

[0009] The analysis of physical properties of materials by processing spectral data (1D) and images (2D) obtained by microscopy has been developed for this purpose.

[0010] The images and spectra are for example obtained by high-resolution transmission electron microscopy (HRTEM and HRSTEM and 4D-STEM), X-ray fluorescence microscopy (HRSTEM-EDX), electron energy loss microscopy (HRSTEM-EELS and HRSTEM-VEELS), energy-filtered transmission electron microscopy (EF-HRTEM, EF-HRSTEM), atomic force microscopy (AFM), scanning tunneling microscopy (STM), atom probe tomography or electron tomography for example. The microscopy images obtained are atomic resolution images, i.e. with a spatial resolution of less than or equal to 2 nanometers.

[0011] Spectral and image data sets each representing at least a portion of the observed sample are obtained by any type of microscopy characterization generating an N-dimensional data block (or multidimensional data), N being an integer greater than or equal to 2, also called a "datacube". This block comprises drops, also called spots, or "spikes" or "blobs" in English, which stand out against a homogeneous background (for example, light drops or spots on a dark homogeneous background). These drops are representative of structural characteristics of the material of the observed sample, for example, an alignment of atoms along the observation direction. The data block may also contain noise.

[0012] For example, when observing crystals, drops arranged in a regular pattern typically represent the crystal lattice. In two dimensions, drops are also called spots. In one dimension, drops are points.

[0013] Mathematically, a drop is a simply connected component of a discrete topological space, in the sense of general topology. This means that any loop traced in a drop can be reduced by homotopy to a point. Physically, a drop represents, for example, the electrical signal produced by the pixels of a matrix detector following the impact of a particle (electron, photon, ion, fermions, boson, etc.). The particle is always much smaller than the pixel, so the drop can always be reduced to a point by homotopy, in accordance with the mathematical definition of the drop. There is therefore a precise agreement between the physical definition and the mathematical definition of the drop.

[0014] In the state of the art, methods are known for measuring dislocations in materials observed by microscopy, using the Radon transformation. One of the difficulties is to locate this type of defect with good precision.

[0015] There is therefore a need to improve the precision of localization of the most difficult defects to detect, in particular dislocations or stacking faults in particular, in observed samples, in particular samples of crystalline material.

[0016] To this end, the invention proposes, according to one aspect, a method for processing multidimensional data representative of a sample of crystalline material for the localization of defects in said sample of crystalline material, comprising an acquisition of at least one image forming a block of input data, the or each image of said block of input data being representative of a part of said sample, said block of input data being represented in an N-dimensional space, N being greater than or equal to two, each data item of said block corresponding to a point in the N-dimensional space. This method comprises steps of: - transforming said block of input data to obtain a block of drops representative of an atomic structure of the observed sample, - into at least one subset of points of said block of drops, • calculation of a structure tensor as a function of the directional gradient values, around each point P of said subset of the droplet block, • calculation of one or more characteristic values ​​of the structure tensor, • formation of a signature tensor associated with the subset of points P of the droplet block, the signature tensor having as components said characteristic values ​​of the structure tensor; - obtaining a block of signatures corresponding to said block of drops, said block of signatures bringing together the signature tensors associated with the points P of said block of drops; - detection and localization of a fault by applying a machine learning classification method to the signature block.

[0017] Advantageously, the use of characteristics of the structure tensor makes it possible to obtain a set of signature tensors, and subsequently, the detection and localization of these defects, in particular dislocations and stacking faults, with a precision impossible to obtain with the methods of the prior art.

[0018] The method for processing multidimensional data for fault location according to the invention may also have one or more of the characteristics below, taken independently or in any technically feasible combinations: possible.

[0019] The characteristic values ​​of a structure tensor are chosen from gradients in each direction, the eigenvalues ​​of the structure tensor, the energy of the structure tensor, the coherence of the structure tensor, the orientation of the structure tensor, the Harris index of the structure tensor, the Harris-Laplace index of the structure tensor.

[0020] The classification method applied is a method developed by supervised, semi-supervised or autonomous machine learning after prior learning on signature blocks calculated on simulated defects.

[0021] The neighborhood is defined by a number of points between half and ten times the median distance between neighboring drops.

[0022] The method comprises a calculation of the eigenvalues ​​of said matrix, and a determination of the minimum eigenvalue, of the maximum eigenvalue, said minimum and maximum eigenvalues ​​being stored.

[0023] The method further comprises an energy calculation, the energy being equal to the sum of the absolute values ​​of said eigenvalues, a coherence calculation, the coherence being equal to the difference between the maximum eigenvalue and the minimum eigenvalue divided by the sum of the maximum eigenvalue and the minimum eigenvalue, said characteristics of the structure tensor comprising the minimum eigenvalue, the maximum eigenvalue, the energy and the coherence.

[0024] The method further comprises an orientation calculation as a function of the structure tensor.

[0025] The method further comprises a calculation of the Harris index and / or the Harris-Laplace index by applying a calculation operator to the values ​​of the structure tensor.

[0026] The invention also relates to a computer program comprising software instructions which, when executed by a computer, implement a method of processing multidimensional data for the location of defects as defined above.

[0027] The invention also relates to a device for processing multidimensional data representative of a sample of crystalline material for the location of defect(s) in the sample of crystalline material, configured to acquire at least one microscopy image forming a block of input data, the or each image of said block being representative of a part of the observed sample, said block of input data being represented in an N-dimensional space, N being greater than or equal to two, each data item of said block corresponding to a point in the N-dimensional space, the device comprising a calculation processor configured to implement: a module for transforming said input data block to obtain a block of drops representative of an atomic structure of the observed sample, - apply, in at least a subset of points of said droplet block, • a module for calculating a structure tensor as a function of the directional gradient values, around each point P of said subset of the droplet block, • a module for calculating one or more characteristic values ​​of the structure tensor, • a module for forming a signature tensor associated with the subset of points P of the droplet block, the signature tensor having as components said characteristic values ​​of the structure tensor,

[0028] - a module for obtaining a block of signatures corresponding to said block of drops, said block of signatures bringing together the tensors of signatures associated with the points P of said block of drops;

[0029] - a module for detecting and locating a fault by applying a machine learning classification method to said signature block.

[0030] The invention will appear more clearly on reading the description which follows, given solely by way of non-limiting example, and made with reference to the drawings in which:

[0031] [Fig-1] [Fig.l] is a block diagram of a multi-channel data inspection system microscopy data comprising a device for processing multidimensional microscopy data according to one embodiment;

[0032] [Fig.2] [Fig.2] is a synopsis of the main stages of a treatment process multidimensional data for fault localization according to one embodiment;

[0033] [Fig.3] [Fig.3] is an example of an input data tile and a drop tile in two dimensions;

[0034] [Fig.4] [Fig.4] illustrates characteristics of structure tensors calculated for the droplet pavement of [Fig.3].

[0035] [Fig.l] schematically illustrates a system for detecting and locating defects 2, in particular dislocations, in a sample of crystalline material from multidimensional data representative of the sample, acquired by a characterization machine 4.

[0036] Any microscopy technique suitable for the characterization of such a sample is applicable.

[0037] In one embodiment, the characterization machine 4 is a transmission electron microscope (TEM) which makes it possible to acquire images of a sample. composed of one or more materials, including for example crystalline materials.

[0038] The electron microscope 4 makes it possible to simultaneously acquire electron diffraction images, EELS spectra (for “Electron Energy Loss Spectroscopy”), EDX spectra (for “Energy Dispersive X-Ray Analysis”) of the signals from different sensors (BF for “Bright field”, DF for “Dark field”, DPC for “Differential phase contrast” or any other suitable or customized sensor), for each point of the sample in scanning microscopy mode called STEM mode (for “Scanning Transmission Electron Microscope”). Another possible acquisition mode is the TEM mode which makes it possible to obtain a global image without having to scan the electron beam.

[0039] In another embodiment, the characterization machine 4 is an atomic force microscope or a scanning tunneling microscope.

[0040] Each acquired spectrum or image is represented in the form of a digital image, comprising points or pixels, each image being representative of at least part of the observed sample.

[0041] All the acquired spectra and images form an N-dimensional data block, N being a natural integer greater than or equal to 2. Such a data block is also called a “datacube”.

[0042] In one embodiment, multiple spectra or images are acquired over time, showing an evolution of the sample during analysis. In this embodiment, time is an additional dimension of the data block.

[0043] Additionally, another dimension of the data block is the microscope focus, electron energy, an angle (of the sample, electron collection, electron convergence), or any other microscope adjustment parameter that may vary in a controlled manner during the measurement.

[0044] Incomplete data can possibly be extrapolated from neighboring values ​​if necessary, for example to correct any imperfections encountered during data acquisition. As a general rule, measurements at the atomic scale are very often tainted by defects because a simple acoustic or electronic noise can sometimes disturb them.

[0045] In a multidimensional data block of a crystal structure sample, the crystal lattice usually forms a periodic network of observable drops arranged in a regular pattern that represents the regular crystal lattice.

[0046] The structural defects or irregularities of such a structure also form drops, but which are for example irregularly arranged, and sometimes smaller than the drops belonging to the regular pattern of the crystal.

[0047] A dislocation-type defect corresponds to a sliding of a part of the crystal along a plane. In N dimensions, this corresponds to a sliding of a part of the data along a hyperplane. Local symmetry breaking can lead to a large number of different defects listed in the literature.

[0048] As already indicated above, mathematically, a drop is a simply connected component of a discrete topological space, in the sense of general topology.

[0049] The values ​​of the pixels belonging to a drop are distinguished from an image background.

[0050] A block of multidimensional data of a sample of a physical structure, e.g. a crystalline material, observed is transmitted to a device 6 for processing multidimensional microscopy data for the localization of defects in the observed sample.

[0051] For example, the transmission is carried out by a wired connection or by a wireless connection (optical, radio, or other).

[0052] The processing device 6 is, in one embodiment, a programmable electronic device, e.g. a computer.

[0053] The device 6 comprises a processor 8 (CPU or GPU) associated with an electronic memory 10. Optionally, the device 6 comprises a human-machine interface 12, comprising in particular a data display screen. In addition, the device 6 comprises or is connected to a storage memory 14. The elements 8, 10, 12, 14 of the device 6 are adapted to communicate via a communication bus 16.

[0054] The processor 8 is configured to execute modules 18, 20, 22, 24, 26, 28, stored in the electronic memory 10, to implement a method for processing multidimensional data representative of the observed sample.

[0055] The module 18 is a module for obtaining a multidimensional data block to be processed, configured to obtain a data block comprising at least one input image, from a data block obtained by the characterization machine 4.

[0056] For example, the obtaining module 18 is configured to obtain the data block to be processed (or input data block) from an electronic memory, where this data has been stored after acquisition.

[0057] Module 20 is a module for transforming the input block into a block of drops (or simply connected components) representative of the atomic structure of the observed sample.

[0058] For example, module 20 implements data filtering, normalization and thresholding or segmentation by machine learning as described below.

[0059] Module 22 is a module for calculating a structure tensor as a function of the directional gradient values, around each point P of a subset of the droplet block. The structure tensor is calculated as a function of the directional gradient values tional, each direction corresponding to a dimension of space, in a neighborhood of the points, the structure tensor being associated with each point of the subset of points.

[0060] For example, the chosen subset of points contains all the points P of the droplet tile for which the calculation of the structure tensor is possible, depending on the size of the neighborhood around each point.

[0061] Module 24 is a module for calculating one or more characteristics of the structure tensor, making it possible to form a signature tensor associated with the subset of points of the processed droplet block.

[0062] Module 26 is a module for obtaining a signature block bringing together the signature tensors associated with the points of the drop block.

[0063] Module 28 is a module for detecting and locating a defect by applying a machine learning classification method to the signature block, supervised or semi-supervised or autonomous after learning on signature blocks calculated on simulated dislocations.

[0064] The parameters 30 of the classification method, obtained by machine learning, are stored, for example in the electronic storage memory 14.

[0065] Semi-supervised or autonomous machine learning is applied in one step prior learning, by an electronic computing device of the same type, on simulated data blocks, the defects of which are also simulated and therefore the theoretical characteristics of which are known. The structure tensor characteristics are calculated on such simulated data blocks, so as to train the classification method, and in particular to obtain by machine learning (possibly of the “deep learning” type) the values ​​of the parameters of the method to be applied.

[0066] For example, machine learning is performed with a fast random forest algorithm, described in the article by Léo Breiman (2001). “Random Forests”, published in Machine Learning. 45(1):5-32).

[0067] Alternatively, other machine learning classification algorithms may be implemented, for example Bayesian type, or function-based, or rule-based, or a combination of such algorithms, or based on the entire range of tools used in machine learning classification.

[0068] According to another variant, the chosen classification method implements a neural network, trained by supervised learning, for example a convolutional neural network CNN (for “Convolutional Neural Networks”) or “deep learning” type learning.

[0069] In one embodiment, the modules 18, 20, 22, 24, 26, 28 are each produced in the form of software, and form a computer program, comprising software instructions which, when executed by a computer, implement implements a multidimensional data processing method for fault localization as described in more detail below.

[0070] In a variant not shown, the modules 18, 20, 22, 24, 26, 28 are each produced in the form of a programmable logic component, such as an FPGA (Field Programmable Gate Array), a GPU (graphics processor) or a GPGPU (General-purpose processing on graphics processing), or in the form of a dedicated integrated circuit, such as an ASIC (Application Specific Integrated Circuit).

[0071] The multidimensional data processing software for fault location is further capable of being recorded, in the form of a computer program comprising software instructions, on a computer-readable medium, not shown. The computer-readable medium is, for example, a medium capable of storing the electronic instructions and of being coupled to a bus of a computer system. By way of example, the readable medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example EPROM, EEPROM, FLASH, SD, NVRAM, RRAM, PCRAM, 2D NAND, 3D NAND, SLC NAND, MLC NAND, TLC NAND, V-NAND, QLC), a magnetic card or an optical card, an SSD disk. [Fig.2] is a synopsis of the main steps of the multidimensional microscopy data processing process for the localization of defect(s) in a sample of crystalline material.

[0072] The method comprises a step 32 of obtaining a block of input data to be processed, comprising at least one image, representing a sample of crystalline material.

[0073] The data block comprises drops standing out from a homogeneous background, for example light drops on a dark background in HRSTEM-HAADF, or HRTEM microscopy, representative of regular patterns of the mesh and defects or irregularities in the observed sample.

[0074] Generally, the input data block is an N-dimensional block, N greater than or equal to 2.

[0075] The data in the tile are numeric values, each value being associated with a point in N-dimensional space. Such a point is also called a pixel when N=2. Each point has an associated coordinate in each dimension, usually represented by an index.

[0076] In one embodiment, N=2, the input data block then being an LxW matrix (i.e. L rows and W columns), composed of pixel values, each pixel having respective coordinates (x,y), x being for example a row index and y a column index.

[0077] It is considered that the input data block is homogeneous, i.e. corresponds to a homogeneous crystalline material, having a homogeneous structure.

[0078] Optionally, if this is not the case, a division of the input data block into homogeneous input data blocks is implemented. Each homogeneous input data block comprises homogeneous data according to a homogeneity criterion relating to at least one material of the observed sample. The homogeneity criterion is a structural homogeneity, with reference to the similarity of the representative domains. Such a division is for example carried out by machine learning, for example by supervised, semi-supervised segmentation or by autonomous classification after learning (for example deep).

[0079] In general, preferably, at least 3 to the power of N neighboring drops are used to define the size of a representative domain of the structure, where N represents the dimension of the droplet tile.

[0080] For an image, at least G=9 neighboring drops are therefore required to define a domain representative of the structure. A component is said to be homogeneous when the Fourier transforms of all its representative domains are all similar. As a reminder, there are many classical methods for mathematically quantifying the similarity between two images. The Fourier transform gives an image of the structural order, similar to the crystalline order revealed by an electron diffraction image.

[0081] According to variants, the homogeneity criterion is a criterion of homogeneity of data intensity values ​​and / or dielectric permittivity and / or electromagnetic properties, for example measured in situ during the acquisition of the data block. The VEELS measurements make it possible in particular to obtain the local dielectric permittivity, as shown in the literature. This criterion of structural homogeneity or physical properties therefore makes it possible to carry out this partitioning of the datacube into areas of interest.

[0082] Step 32 is followed by a step 34 of transforming said input data block to obtain a block of drops representative of an atomic structure of the observed sample.

[0083] As already indicated above, the drops are simply connected components of the N-dimensional discrete topological space.

[0084] In a first embodiment, step 34 implements a filtering of the component sub-tile, optionally a normalization, then a segmentation.

[0085] For example, step 34 implements convolution filtering with an HBSG kernel (for “Half Bail Savitzky-Golay”) as described in patent FR3 118 256. This filtering, for example, uses a square HBSG kernel of size corresponding to the median of the distances between neighboring drops (here 21 pixels), with smoothing of order 2 or 3, then local normalization of the gray levels. Local normalization consists of dividing the image by the local variance, obtained by calculating the square root of the convolution by a Gaussian filter of the square of the image. Generally, the size of the Gaussian filter used to calculate the local variance is ideally close to the median of the distances between neighboring drops, .and Finally a segmentation is carried out, for example by applying an adaptive thresholding of Bemsen described in "Dynamic Thresholding of Grey-Level Images", Proc, of the 8th Int. Conf. on Pattern Recognition, 1986, with a radius close to half of the median of the distances between neighboring drops (here 10 pixels).

[0086] A second segmentation variant that is more precise than Bernsen's method implements machine learning as described in more detail below. The machine learning algorithm makes it possible to distinguish two classes, respectively the "drop" class (or "spots" in English) and the "background" class.

[0087] For example, machine learning is performed with a fast random forest algorithm, described in the article by Léo Breiman (2001). “Random Forests”, published in Machine Learning. 45(1):5-32).

[0088] Alternatively, other machine learning classification algorithms may be implemented, for example Bayesian type, or function-based, or rule-based, or a combination of such algorithms, or based on the entire range of tools used in machine learning classification such as deep learning for example.

[0089] Typically, droplets represent less than 20% of the image, and their size is less than 1 nm, after calibration to the actual sizes in the sample. Simulations of the theoretical images allow the expected distribution of droplets in the image to be determined in order to verify that the segmentation is correct. A segmentation is correct if the number and position of droplets approach the expected theoretical values ​​beyond a given uncertainty threshold.

[0090] Optionally also, convolution filtering to fill any holes inside the obtained drops is implemented. For example, a filter as described in the patent published under number FR 3 118 256 is implemented, and / or a “fill holes” type algorithm (K.-J. Oh, S. Yea, and Y.-S. Ho, “Hole filling method using depth based inpainting for view synthesis in free viewpoint television and 3-D video,” in Proc. Picture Coding Symp., Chicago, IL, May 2009.). Again, the objective is to approach the ideal theoretical shape, which is free of holes.

[0091] Another option to get even closer to the theoretical distribution of the drops consists of carrying out a theoretical convolution, i.e. a masking of the diffractogram of the droplet block with a theoretical diffractogram adjusted as best as possible to the experimental diffractogram. This adjustment of the theory with respect to the experiment consists of varying the theoretical parameters of lengths (a,b,c) and angles (a,[3,y) defining the crystal mesh in order to minimize the mean square deviation between the experimental diffractogram and the theoretical diffractogram. Another variant of diffractogram segmentation uses machine learning that makes it possible to distinguish two classes, respectively the “diffraction peak” class and the “background” class.

[0092] For example, machine learning is performed with a fast random forest algorithm, described in the article by Léo Breiman (2001). “Random Forests”, published in Machine Learning. 45(1):5-32).

[0093] Alternatively, other machine learning classification algorithms may be implemented, for example Bayesian type, or function-based, or rule-based, or a combination of such algorithms, or based on the entire range of tools used in machine learning classification such as "deep learning" for example. This machine learning segmentation makes it possible to retain only the significant frequency peaks and to selectively eliminate unwanted information from the diffractogram.

[0094] The most detailed option for extracting the droplet distribution is to mask the component diffractogram with a mask that is the maximum of (A) the adjusted theoretical diffractogram normalized between 0 and 1, and (B) the machine learning segmentation result normalized between 0 and 1 of the Fourier transform of the droplet tile.

[0095] It is the atomistic simulation of the “perfect” theoretical images which makes it possible to verify that the image processing parameters make it possible to converge towards the perfect theoretical form, for each crystal, each orientation and each set of instrumental parameters.

[0096] The method then comprises the following steps, repeated for each block of drops.

[0097] In at least one subset of points P of the droplet block, each point P having coordinates Xp=(xl_p,x2_p,.. .,xN_p), the method implements a calculation 36 of a structure tensor as a function of the directional gradient values, each direction corresponding to a dimension of the N-dimensional representation space.

[0098] In differential geometry, the structure tensor or second moment matrix is ​​classically defined by a matrix derived from the gradient of a function, describing the distribution of the gradient in a specified neighborhood around a point. The gradient is oriented towards the direction of the greatest variation of the scalar field. The structure tensor is used in image processing and allows the analysis of the local anisotropy around a given point by estimating the predominant directions of the gradient in the neighborhood of this point.

[0099] For N=2, the structure tensor field of an image is generally defined as the field of local covariance matrices of the partial first derivatives of this image. It is constructed from the gradient fields previously estimated by linear convolution, which leads to the following mathematical formula: [Math.l]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109] Where / v is the partial derivative in the first dimension, j Xn is the partial derivative in the second dimension, and w() is an isotropic observation window, centered at P, with a generally Gaussian weighting kernel. The symbol denotes the convolution operation. This spatial smoothing by convolution makes the structure tensor more robust to noise. The choice of the spatial extension of this smoothing is therefore decisive to obtain the best compromise between precision and robustness to noise. The structure tensor field generally has the same size as the data block, except for the edges where calculation by convolution is not possible. On the other hand, it is a field with matrix values. Thus, in one embodiment, the subset of points P is for example the subset of points which contains all the points P of the droplet block for which the calculation of the structure tensor is possible (i.e. calculation by possible convolution). In 3 dimensions, the formula [MATH 1] becomes: [Math.2] The generalization of the formula [MATH 2] to N dimensions, with N>3, is analogous, by calculating the field of local covariance matrices of the partial first derivatives in the corresponding hyperspaces. Subsequently, the formulas are given for N=2, the generalization to N>2 being within the reach of those skilled in the art. For example, in one embodiment, the partial derivatives are obtained by calculating directional gradients, according to each of the dimensions of the space. The structure tensor is defined by the matrix of [MATH 1], applied with the values ​​of the directional gradients in each direction. The neighborhood of points taken into consideration is for example defined by a number No of neighboring points around the point P considered in each direction. For example, No is chosen based on the minimum distance between two drops neighbors for a given crystal lattice structure.

[0110] In one embodiment, No is between half and ten times the median distance between neighboring drops.

[0111] The method also comprises a step 38 of calculating one or more characteristic values ​​of the structure tensor.

[0112] In one embodiment, this step of calculating characteristic values ​​of the structure tensor comprises several sub-steps.

[0113] Thus, the calculation step 38 comprises a sub-step 40 of calculating the eigenvalues ​​of the structure tensor matrix associated with the point P considered.

[0114] In the case where N=2, the structure tensor at each point P considered is a 2x2 matrix, which includes two associated eigenvalues, which are calculated analytically.

[0115] The eigenvalues ​​of the structure tensor associated with the point P considered are noted respectively Xi(p) and X2(p).

[0116] The eigenvalues ​​make it possible to obtain information on the distribution of the gradient in the window w, and to discriminate between regions with directional gradient, uniform regions or regions with rotational symmetry.

[0117] The eigenvalues ​​are stored in substep 40.

[0118] In addition, the method comprises a sub-step 42 of calculating an energy value associated with the structure tensor, by applying the following formula:

[0119] [Math.3] E p — p / p)! + |22(p)|

[0120] The energy of a structure tensor, defined as the sum of its eigenvalues, characterizes the dynamics. The energy of the tensor at a point expresses the local contrast in the vicinity of this point. (J. Angulo. Structure Tensor Image Filtering using Riemannian L1 and Léo Center-of-mass. Image Analysis and Stereology, vol. 33, no. 2, pp. 95-105,2014.)

[0121] If the energy is close to 0, the point P belongs to a homogeneous region.

[0122] The energy value associated with the point considered is stored.

[0123] The method also comprises a sub-step 44 of calculating the associated coherence, the coherence being defined by the difference between the maximum eigenvalue and the minimum eigenvalue divided by the sum of the maximum eigenvalue and the minimum eigenvalue. The coherence is also called confidence factor or dispersion indicator. The coherence represents a measure of the local anisotropy and characterizes the dispersion of the orientation of the gradient, i.e. the local variability of the geometry of the image.

[0124] Let Xmax(p)=max(Xi(p), X2(p)) be the maximum eigenvalue and Xmin(p)=min(Xi(p),X2(p)) the minimum eigenvalue, the formula for calculating the coherence is:

[0125] [Math.4] Ap)

[0126] The coherence value associated with the point considered is stored.

[0127] The calculation step 38 further comprises a sub-step 46 of calculating an orientation associated with the point considered, from the components of the structure tensor. The method for calculating the orientation is for example described in the following publication: IL Dryden, A. Koloydenko and D. Zhou. Non-Euclidean Statistics for Covariance Matrices, with Applications to Diffusion Tensor Imaging. The Annals of Applied Statistics, vol. 3, no. 3, pp. 1102-1123, 2009.

[0128] The orientation value associated with the point considered is stored.

[0129] The calculation step 38 further comprises a sub-step 48 of calculating a Harris index value at point P and / or calculating a Harris-Laplace index value at point P. The Harris index, known in the field of image processing, is used for detecting corners in digital images. Similarly, the calculation of the Harris-Laplace index is known in image processing. The Harris index and / or Harris-Laplace values ​​are stored.

[0130] Thus, at the end of step 38 of calculating characteristic values ​​of the structure tensor, a signature tensor associated with the point P considered is formed (step 50), the components of which are characteristic values ​​previously calculated.

[0131] In one embodiment, the characteristics of the structure tensor retained in the signature tensor comprise: the gradient values ​​along each of the dimensions, the eigenvalues ​​of the structure tensor, the energy value, the coherence value, the orientation value, the Harris index value, the Harris-Laplace index values.

[0132] According to variants, the characteristics of the structure tensor retained in the signature tensor comprise a subset of one or more characteristics mentioned above.

[0133] It should be noted that in such a variant, certain calculation sub-steps are not carried out, for example if the energy value is not retained among the components of the signature tensor, the sub-step of calculating and storing the energy value is not carried out.

[0134] Steps 38 to 50 are repeated for each point of the subset of points considered.

[0135] Thus, for each droplet block considered, an associated signature tensor is obtained and stored.

[0136] After processing all the points considered, the method comprises a step 52 of obtaining a block of signatures, associated with the block of drops, bringing together the signature tensors associated with the points P of the droplet block considered.

[0137] This signature block is provided as input to a step 54 of detecting and locating fault(s).

[0138] Step 54 implements a method for classifying the signature block into one or more defect categories. The classification method is supervised, semi-supervised or autonomous by prior machine learning carried out from the most exhaustive possible family of defects originating from simulated data blocks, for theoretical crystal structures.

[0139] For example, machine learning is performed with a fast random forest algorithm, described in the article by Léo Breiman (2001). "Random Forests", published in Machine Learning. 45(1):5-32).

[0140] Alternatively, other machine learning classification algorithms may be implemented, for example Bayesian type, or function-based, or rule-based, or a combination of such algorithms, or based on the entire range of tools used in machine learning classification.

[0141] According to another variant, the chosen classification method implements a neural network, trained by supervised learning, for example a convolutional neural network CNN (for “Convolutional Neural Networks”). Learning can also be carried out by deep learning, but this variant is more expensive in terms of computation time.

[0142] Thus, from the signature block, where applicable, a type (or class) of fault and an associated location are obtained as output.

[0143] In the event of fault detection, information relating to the presence of the fault and its location is provided, for example on a map which is displayed on a human-machine interface for an operator.

[0144] Alternatively or additionally, the information relating to the presence of a defect and its location is provided to a control software application, for subsequent action, for example, raising an alert. This makes it possible, for example, to eliminate materials with defects before their use in a manufacturing process for devices based on such materials. This information is also valuable for establishing links between the material production processes and the resulting defect, in a logic of sequential optimization.

[0145] By way of example, [Fig.3] illustrates, by data obtained by simulation, an input block for N=2, i.e. a digital image block 60 and the corresponding drop block 62, which is also a digital image of drops 62, after application of the transformation step 34.

[0146] The droplet block 62 includes a dislocation located in the center, which is particularly difficult to detect (screw dislocation). No state-of-the-art method made it possible to detect it beforehand.

[0147] [Fig.4] illustrates an image 64 of the energy values ​​64 obtained by the calculation of the calculation sub-step 42 on all the points of the digital image of drops 62; the coherence values, combined into a normalized coherence image 66, obtained by the calculation of the calculation sub-step 44 on all the points of the digital image of drops 62; the orientation values, combined into a normalized orientation image 68, obtained by the calculation of the calculation sub-step 46 on all the points of the digital image of drops 62.

[0148] Thus, it is clear that the energy, coherence and orientation characteristics of the structure tensors indicate the presence of a non-homogeneous structure (i.e. here a dislocation), in a well-defined spatial region, and make it possible to locate the dislocation, even if it is almost invisible in the initial image.

[0149] Advantageously, the calculation of a structure tensor and the characteristics of the structure tensor is rapid and not very complex, which makes it possible to accelerate the calculations and to reduce the amount of computational resources required. This calculation advantageously makes it possible to make visible defects hitherto undetected by the usual techniques.

[0150] Advantageously, the accumulation of several characteristics in a signature block allows precise characterization of dislocations of any type (screw; corner, mixed, stacking faults, etc.).

[0151] According to variants, it is also possible to replace the gradient calculation with higher order differentials, for example a linear combination of partial derivatives of order n, which is more complex from a computational point of view, but makes it possible to characterize fluctuations of more complex structures.

[0152] Advantageously, very good spatial precision of dislocation localization is obtained by the proposed method, typically picometric. The method has exceptional sensitivity and applies to all defects which induce local deformation of the crystalline structure.

Claims

Claims

1. Method for processing multidimensional data representative of a sample of crystalline material for the location of defects in said sample of crystalline material, comprising an acquisition of at least one image forming a block of input data, the or each image of said block of input data being representative of a part of said sample, said block of input data being represented in an N-dimensional space, N being greater than or equal to two, each data item of said block corresponding to a point of the N-dimensional space, the method being characterized in that it comprises steps of: - transformation (34) of said block of input data to obtain a block of drops representative of an atomic structure of the observed sample, - into at least one subset of points of said block of drops, • calculation (36) of a structure tensor as a function of the directional gradient values,around each point P of said subset of the droplet block, • calculation (38) of one or more characteristic values ​​of the structure tensor, • formation (50) of a signature tensor associated with the subset of points P of the droplet block, the signature tensor having as components said characteristic values ​​of the structure tensor; - obtaining (52) a signature block corresponding to said droplet block, said signature block bringing together the signature tensors associated with the points P of said droplet block; - detection and localization (54) of a defect by applying a machine learning classification method to said signature block.,

2. A method according to claim 1, wherein said characteristic values ​​of a structure tensor are chosen from gradients in each direction, the eigenvalues ​​of the structure tensor, the energy of the structure tensor, the coherence of the structure tensor, the orientation of the structure tensor, the Harris index of the structure tensor, the Harris-Laplace index of the structure tensor.

3. Method according to one of claims 1 or 2, in which the classification method applied is a method developed by supervised, semi-supervised or autonomous machine learning after prior learning on signature blocks calculated on simulated defects.

4. Method according to one of claims 1 to 3, in which said neighborhood is defined by a number of points between half and ten times the median distance between neighboring drops.

5. Method according to any one of claims 1 to 4, comprising a calculation (40) of the eigenvalues ​​of said matrix, and a determination of the minimum eigenvalue, of the maximum eigenvalue, said minimum and maximum eigenvalues ​​being stored.

6. The method of claim 5, further comprising an energy calculation (42), the energy being equal to the sum of the absolute values ​​of said eigenvalues, a coherence calculation (44), the coherence being equal to the difference between the maximum eigenvalue and the minimum eigenvalue divided by the sum of the maximum eigenvalue and the minimum eigenvalue, said characteristics of the structure tensor comprising the minimum eigenvalue, the maximum eigenvalue, the energy and the coherence.

7. Method according to one of claims 1 to 6, further comprising a calculation (46) of orientation as a function of the structure tensor.

8. Method according to any one of claims 1 to 7, further comprising a calculation (48) of Harris index and / or Harris-Laplace index by applying a calculation operator to the values ​​of the structure tensor.

9. A computer program comprising software instructions which, when executed by a programmable electronic device, implement a method of processing multidimensional data representative of a sample of crystalline material for the location of defects in said sample of crystalline material according to claims 1 to 8.

10. Device for processing multidimensional data representative of a sample of crystalline material for the location of defect(s) in the sample of crystalline material, configured to acquire at least one microscopy image forming a block of input data, the or each image of said block being representative of a part of the observed sample, said block of input data being represented in an N-dimensional space, N being greater than or equal to two, each data item of said block corresponding to a point in the N-dimensional space, the device comprising a calculation processor configured to implement: - a module (20) for transforming said input data block to obtain a block of drops representative of an atomic structure of the observed sample, - apply, in at least a subset of points of said droplet block, • a module (22) for calculating a structure tensor as a function of the directional gradient values, around each point P of said subset of the droplet block, • a module (24) for calculating one or more characteristic values ​​of the structure tensor, • a module (24) for forming a signature tensor associated with the subset of points P of the droplet block, the signature tensor having as components said characteristic values ​​of the structure tensor, - a module (26) for obtaining a block of signatures corresponding to said block of drops, said block of signatures bringing together the tensors of signatures associated with the points P of said block of drops; - a module (28) for detecting and locating a fault by applying a machine learning classification method to said signature block.