Method for determining additional training data for training a constraint field classifier, and associated electronic devices
By forming pure subsets of stress fields and using convex combinations to generate additional training data, the method addresses the computational challenges of simulating turbine blade deformation and aging, enhancing prediction accuracy and reducing simulation time.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- SAFRAN SA
- Filing Date
- 2020-12-16
- Publication Date
- 2026-04-24
AI Technical Summary
Numerical simulations of mechanical parts, particularly turbine blade deformation and aging, require significant computing resources and time due to the complexity of temperature fields, making it difficult to predict long-term mechanical behavior accurately.
A method for determining additional training data using convex combinations of basic stress fields to form pure subsets, which are then used to generate additional training data for a constraint field classifier, reducing computation time while maintaining accuracy.
The method significantly increases the number of reliable training data points without increasing computation time, improving the precision of mechanical response predictions for turbine blades and other highly stressed components.
Smart Images

Figure 00000045_0000 
Figure 00000045_0001 
Figure 00000045_0002
Abstract
Description
Title of the invention: Method for determining additional training data for training a strain field classifier, and associated electronic devices. TECHNICAL FIELD OF THE INVENTION
[0001] In general, the invention relates to the field of numerical simulation of the mechanical behavior of mechanical parts, for example, their deformation and aging. It relates in particular to the application of artificial intelligence techniques during a preliminary simplification phase of the numerical simulation to be performed. The invention specifically relates to the determination of training data for training a stress field classifier, for example, a temperature field, this classifier enabling the identification, from a pre-established list, of a simplified mechanical behavior model well-suited to the stress field under consideration. TECHNOLOGICAL BACKGROUND OF THE INVENTION
[0002] Numerical simulation using finite element methods, which simulates the deformations undergone by a mechanical part subjected to various stresses, has made significant progress in recent years. Such simulations make it possible, in particular, to predict how a given mechanical part will age over a long period of use (for example, several years), which could not have been determined through direct testing or measurement. More generally, these simulations make the design and prototyping phase of a mechanical system faster and more efficient.
[0003] But some situations remain difficult to simulate, because they require a very large amount of computation time.
[0004] This is particularly relevant for characterizing the aging of a turbine blade in a turbomachine, for example, a high-pressure turbine located just downstream of the combustion chamber. In this case, a high, and generally inhomogeneous, temperature is exerted on the turbine blades, in addition to the pressure field and centrifugal force experienced by these blades.
[0005] For a given temperature field, exerted on this blade, simulating the deformation and aging of this or that element of the blade, for several cycles corresponding to several flights of an aircraft, requires very significant computing resources.
[0006] Furthermore, the temperature field exerted on such a turbine blade is not actually perfectly known. In order to be able to precisely determine a duration during which a blade will remain usable, and an uncertainty associated with this duration, it is necessary to therefore repeat the numerical simulation (already very demanding in itself) several times, for a whole (statistical) distribution of temperature fields.
[0007] With a classical finite element numerical simulation, the computation time or computing resources for such a characterization (number of processors or computing clusters to be made to work in parallel) then becomes prohibitive.
[0008] To reduce the computation time required to perform such simulations, an approach based on grouping different possible inputs (stresses) into clusters and a classification to recommend a reduced model (e.g., a reduced model of mechanical behavior) was proposed in the article "Localized Discrete Empirical Interpolation Method" by B. Peherstorfer et al., SIAM J. Sci. Comput. (2014). More recently, an approach based on such an order reduction of the mechanical problem and on a temperature field classifier was proposed in the following article: "Model order reduction assisted by deep neural networks (ROM-net)" by T. Daniel et al., Adv. Modell. and Simul. in Eng. Sci. (2020) 7:16.
[0009] According to this approach, instead of determining the mechanical deformation at each node of a mesh of the blade element to be characterized, with a number of degrees of freedom (scalars) which would therefore be equal to at least three times the number of nodes of the mesh, the mechanical response of the blade element is studied on the basis of a simplified mechanical behavior model, Mk, parameterized by a smaller number of degrees of freedom.
[0010] Such a mechanical behavior model Mk describes, for example, the deformation of the turbine blade element under consideration as a weighted sum of deformation modes typical of that element (for example, a pure bending mode, a mode corresponding to uniform stretching, another mode corresponding to stretching localized at one or the other end of the turbine blade element in question, etc.). The degrees of freedom, or in other words the parameters of this mechanical behavior model, then correspond, in this weighted sum, to the coefficients that weight the contributions of these different deformation modes.
[0011] The mechanical response of the turbine blade element, subjected to a given temperature field T, is then determined over time for several flight cycles by numerical simulation, based on this simplified model Mk. Thanks to the simplified nature of this model, it is possible to perform a simulation corresponding to a significant period of use, while maintaining a limited computation time.
[0012] Such a mechanical behavior model Mk is sometimes called a "reduced model" in this technical field. It is representative of the mechanical response of the element of a turbine blade (in particular, the deformation of this element) when this element is subjected to a given type of stress (for example, a given temperature field T). It is therefore, in a way, a typical response of the turbine blade element. Moreover, in the following, such a model of mechanical behavior Mk is referred to interchangeably as a "reduced model" or a "typical response".
[0013] According to the approach proposed in the article by T. Daniel et al. mentioned above, before performing the numerical simulation in question, based on the reduced model Mk, this reduced model Mk is chosen according to the temperature field T exerted on the turbine blade element 1'. The reduced model Mk is selected, by a temperature field classifier 2', from among different reduced models Mb ... Mk, ..., MP, that is to say, from among different models of the mechanical behavior of the turbine blade element ([Fig. 1]). The reduced model Mk that is selected by the classifier 2' is, among the reduced models Mb ... Mk, ..., MP, the one that is most representative of a complete expected mechanical response for the turbine blade element 1' when it is subjected to the temperature field T in question.
[0014] The classifier 2', which includes for example a neural network, is trained using artificial intelligence techniques during a training step 20 ([Fig.2]). This training is carried out on the training data (TH>i, Mk), determined beforehand during a training data determination step 10.
[0015] This step 10 comprises a first step 11 of numerical simulation on a stress cycle of the turbine blade element ([Fig. 3]), corresponding, for example, to the stress experienced by this element during the entire flight of an aircraft powered by the turbomachine to which this blade element belongs. During step 11, the mechanical response of the turbine blade element is determined over time for this stress cycle, the blade element being subjected to a given temperature field TH i.
[0016] Such a numerical simulation is not necessarily sufficient to determine the long-term evolution of the mechanical characteristics of the blade element (which would require simulation over many successive cycles), nor its lifespan. However, it does allow for the identification of the type of mechanical behavior of this element when subjected to the temperature field TH i in question.
[0017] This comprehensive, but short-time, numerical simulation is performed for different so-called "high-resolution" temperature fields, TH.i, i=l,...,N. Each of these high-resolution temperature fields yields temperature values '-rJ 1 Hj to the different nodes j, j=l,...,QH of a mesh (so-called "high resolution" mesh) of the blade element.
[0018] This numerical simulation makes it possible to determine "complete" mechanical responses of the blade element, for these different temperature fields TH,i.
[0019] Then, in step 12, for each pair (TH>ib TH>i 2) grouping two of these high-resolution temperature fields, a dissimilarity of mechanical behavior ô; u 2 is determined, or, in other words, an effective “distance” between the mechanical responses obtained for these two temperature fields TH>ib TH>i 2. This distance ô; u 2 is, for example, representative of a difference between the two strain fields obtained respectively for these temperature fields. Different metrics are conceivable for determining this difference: it can be determined as being equal to the square root of the sum of the squares of the deviations at each node of the mesh, or as the sum of the absolute values of the deviations at each node of the mesh, or even as a distance in the sense of Grassmann, as is done in the aforementioned article by T. Daniel et al.
[0020] Next, in step 13, the various high-resolution mechanical responses determined in step 11 are grouped into several mechanical response groups Cb Ck, ..., CP, each grouping mechanical responses that are similar to each other, i.e., that exhibit, pairwise, a low dissimilarity δ; u 2 (a small "distance" between two elements of the same group). This step is generally called "clustering" in English (i.e., "data partitioning"), and the mechanical response groups Cb C2, CP are generally called "clusters".
[0021] Then, in step 14, for each cluster Ck, k=l.. .P, a reduced model Mk representative of the complete mechanical responses of the cluster considered (responses which are relatively similar to each other, by construction), is determined by an order reduction technique.
[0022] The training data ultimately obtained comprises various "labeled" data, each labeled data (TH.i, Mk) comprising: • one of the high-resolution temperature fields TH>i, i=l,.. ,,N, • as well as a label that identifies, among responses- types of blade element Mk , k=l.. .P, the standard response Mk corresponding to the temperature field TH>i considered (this standard response being the standard response associated with the cluster Ck in which is found the complete, high-resolution mechanical response obtained in response to the temperature field TH,i in question).
[0023] The label in question corresponds for example to the index k which identifies the standard response Mk, in the set of standard responses of the dawn element (this label is for example constituted by the index in question).
[0024] This set of labeled data (TH,i, Mk), i=l,.. .,N, is then used for training, i.e. for learning the classifier 2'.
[0025] It is important to note that the temperature fields TH i are grouped into different clusters Ck, each associated with a given typical response, not on the basis of a similarity between temperature fields, but on the basis of a similarity between the mechanical responses corresponding to these different temperature fields.
[0026] This distinction is particularly important because two temperature fields that appear close and similar to each other can in fact lead to very different mechanical responses, the interactions between thermal and mechanical loading being complex and dependent on the geometry of the turbine blade element.
[0027] Thus, two identical temperature fields, except for a very small portion of the blade element, can in fact lead to very different mechanical responses if the area in question is mechanically critical (for example, because it is thin, or subjected to high mechanical stress, or subject to stress concentrations). Conversely, two temperature fields that differ from each other over a large but less critical area, but which, on the other hand, take on the same values at the critical area in question, will probably lead to similar mechanical responses, even though these two temperature fields are apparently dissimilar.
[0028] The training data (TH>i, Mk), built on the basis of a similarity between responses, therefore allows a training of the classifier which is much more precise and reliable than with training data which would be based on a similarity between solicitations.
[0029] However, obtaining the training data (TH>i, Mk) requires significant computation time, since a numerical simulation (admittedly short-time, but on a complete mesh) is performed for each high-resolution temperature field TH,i. In practice, the number N of labeled data finally obtained is therefore limited.
[0030] For a mesh comprising approximately 40,000 nodes, obtaining just 500 labeled data points already requires a very significant computation time. And if one wants to obtain approximately 104 labeled data points, even with a reduced mesh comprising With only a few thousand nodes, it takes more than a day of computing by parallelizing the calculation on about fifty processors.
[0031] In this context, it would therefore be desirable to be able to obtain a larger number of reliable training data without significantly increasing computation time.
[0032] Various techniques for augmenting training data exist, particularly in the field of automatic image analysis. One known technique is based on applying various transformations and deformations to a given image. For example, an image representing a car is stretched, flipped, blurred, cropped, etc., to obtain additional images, which still represent a car.
[0033] But this type of technique is not applicable to the case considered here. Indeed, deforming one of the temperature fields THi would probably lead to a new temperature field giving a mechanical response completely different from that obtained with the initial temperature field TH,i (in other words, the label would not be preserved during such a transformation).
[0034] On the other hand, a difficulty arises from the size of each labeled data point itself. Indeed, to properly construct the Ck clusters and the associated typical responses Mk, given that the mechanical behavior of the blade element is initially unknown, the simulations mentioned above (short-time simulations) must be performed with a fine mesh. This is because no information is available in advance regarding the more or less critical nature of any given area from a thermomechanical point of view. The temperature field THi is therefore a high-resolution temperature field, and the resulting labeled data (TH>i, Mk) is thus large and cumbersome to process. Summary of the invention
[0035] In this context, a method is proposed for determining additional training data for training a constraint field classifier, • the additional training data being determined from the basic training data, • the basic training data comprising basic stress fields applied to a turbine blade element and, for each of these fields, a label that identifies a typical response, representative of an expected mechanical response for that turbine blade element in response to said stress field, among different typical responses of said turbine blade element, • The method includes the following steps: • form at least one pure subset, which groups together basic stress fields associated with the same standard response of the turbine blade element, and which has a convex envelope containing no basic stress field associated with another standard response of said element, • for at least one of said pure subsets, determination of one or more additional training data, each additional training data comprising an additional constraint field, obtained by linear combination of basic constraint fields of the pure subset considered, assigned positive weighting coefficients and normalized to one, the additional constraint field being accompanied by a label which is the same as for the basic constraint fields of said combination.
[0036] The stress fields in question may, for example, be temperature fields or pressure fields exerted on the turbine blade element. For the sake of simplicity, the technical effects obtained using this method are presented below in the case of temperature fields. However, as will be apparent to those skilled in the art, these effects are obtained in the same way in the case of pressure fields, for example.
[0037] The convex hull of a given set of basic temperature fields Ti denotes the hull, that is, the boundary, of the set that groups together all temperature fields T expressed as a linear combination of these basic temperature fields Ti, weighted with positive coefficients and normalized to 1 (that is, such that the sum of these coefficients is equal to 1). In other words, a temperature field lies inside this convex hull if it is equal to such a linear combination of the basic temperature fields of this set. Otherwise, it lies outside this convex hull.
[0038] Figure 10 schematically represents such a convex hull Eb for a pure subset Si of basic temperature fields Ti, each represented by a small black disk. For comparison, a non-pure subset S* is also shown in this figure: its convex hull E* contains a basic temperature field Tj (represented by a cross) which is associated with a standard response Mk that is different from the standard response Mk common to the different elements of the subset S* (elements which are represented by disks). In other words, the convex hull E* of the subset S' contains a temperature field (represented by the cross) which belongs to a cluster C2 different from the cluster Ci to which the elements of the subset S* belong.
[0039] By determining each additional temperature field T;' in the form of a convex linear combination of basic temperature fields T; belonging to the same pure subset, the probability of a labeling error of this additional temperature field T;' is considerably reduced.
[0040] In other words, due to the pure nature of this subset, it is very likely that the standard response, the one best suited to the additional field T'i (the one that best describes the response of the blade element to the field T'i), is indeed the standard response Mk common to the different elements of this subset.
[0041] By way of illustration, for an example embodiment in which 5400 additional temperature fields T'i were determined from 600 basic temperature fields T;, 400 of the additional temperature fields T'i in question were tested to verify that their labeling was correct.
[0042] This verification was carried out by running a short-time numerical simulation for the additional temperature field under consideration, and then determining to which cluster the resulting complete mechanical response actually belongs, as explained in the section on technological background. In this case, the verification showed that none of the 400 temperature fields tested had been mislabeled, which illustrates the reliability of this method for obtaining additional training data.
[0043] It should be noted in this regard that a linear (even convex) combination of basic temperature fields T, belonging to the same cluster Ck, without any particular precautions (without first selecting pure subsets), would, on the other hand, lead to frequent labeling errors for the additional training data. For example, in the very schematic case of [Fig. 10], a convex linear combination of basic temperature fields from the subset S* would produce additional temperature fields, some of which would be labeled k=1, whereas the label k=2 would actually be appropriate.
[0044] Selecting pure subsets beforehand is all the more important in practice since the different clusters Ci, ..., Ck, ..., CP are in fact very intertwined with each other in the space of temperature fields, more so than in [Fig. 10], which is very schematic. As an example, [Fig. 11] represents, in a manner comparable to [Fig. 10] but based on real data, different basic temperature fields belonging to four different clusters, according to a simplified two-dimensional representation (i.e., projected onto a space with two degrees of freedom). Indeed, this figure shows that the temperature fields belonging elements from different clusters are often mixed together, which shows the importance of identifying pure subsets within clusters, which constitute, in a way, bubbles without elements foreign to the cluster in question.
[0045] In other words, in the space of mechanical responses, the different temperature (or, more generally, stress) fields are grouped into clusters that are not intertwined or mixed with each other (given the way the clusters are constructed). But in the space of temperature fields, the different clusters are, in a sense, mixed with each other ([Fig. 11]), and it is precisely in the space of temperature fields that the linear combinations must be performed to obtain new training data, hence the importance of identifying the pure subsets in question.
[0046] A turbine blade element is defined as any component of the blade or the blade attachment system. This could be a turbine blade itself, but also a turbine blade attachment hook. The data augmentation technique described here is particularly useful for simulating the mechanical behavior and aging of turbine blade elements, especially for a high-pressure turbine, due to the inherently demanding and computationally intensive nature of such simulations. However, this technique can also be applied to the characterization of other mechanical parts of a turbomachine, and more generally, to the characterization of any highly stressed mechanical component.
[0047] Furthermore, this technique can also be applied to obtain additional training data of a different type than that mentioned above. In general, it is well suited to obtaining additional training data for training a classifier, in a context where obtaining the training data is computationally expensive, particularly because it requires prior numerical simulations (such as the short-time numerical simulations presented above).
[0048] The method just presented is implemented by a computer or any computer system capable of executing logical operations.
[0049] In addition to the characteristics mentioned above, the method just presented may have one or more of the following complementary characteristics, considered individually or in all technically possible combinations: • said stress exerted on the turbine blade element is a temperature exerted on that element, the basic stress fields and the additional stress fields being respectively basic temperature fields and additional temperature fields; the step of forming one or more pure subsets includes, for each of these subsets, a purity test step of the subset considered, during which it is verified that no basic stress field associated with a standard response different from the standard response common to the basic stress fields of the subset considered, is located inside the convex hull of said subset; Each basic stress field groups a set of values of said stress at different nodes of a given mesh comprising a total number of nodes Q. For each basic stress field, an associated stress subfield is obtained by projecting the basic stress field onto a subspace with a number of degrees of freedom less than Q, and during the purity test step: • a projection test is performed, during which it is determined whether, for the stress subfields associated with the basic stress fields of the subset to be tested, a convex hull of the set grouping these stress subfields contains no stress subfield associated with one of the basic stress fields for which the standard response is different from the standard response common to the basic stress fields of the subset to be tested, • and, when the result of this projection test is positive, it is determined that the subset of stress fields considered is pure; during the purity testing stage, • As long as the result of said projection test is negative, this projection test is executed again, the subspace onto which said basic constraint fields are projected being different from one execution of the projection test to another, • As soon as the result of the projection test is positive, it is determined that the subset of stress fields to be tested is pure, • and, when after a given number of successive executions of the projection test, the projection test has still not given a positive result, it is determined that the subset to be tested is not pure; The step of forming said at least one pure subset includes the following steps: • among the basic constraint fields that are associated with the same standard response of the blade element, selection of one or more basic constraint fields, each constituting a seed for the subsequent growth of the pure subset(s), • for each of said seeds, formation of a subset, initially containing said seed, which is then grown: • by performing an addition step, during which one of the basic constraint fields is added to the considered subset, the added constraint field being associated with the same standard response as the other constraint fields of said subset, then • by performing said purity test step, for the subset thus augmented, • the addition step and the purity test step being executed again, as long as the subset remains pure; when the purity test step shows that the subset is no longer pure, the last basic constraint field to have been added to the subset is removed from the subset to obtain the final pure subset; During the step of selecting one or more seeds, the seeds are selected from among the basic stress fields which are associated with the same standard response of the blade element, and which exhibit, with respect to the different basic stress fields associated with the other standard responses, a dissimilarity of mechanical response greater than a given threshold; during the step of selecting one or more seeds, each seed is selected from the basic stress fields which are associated with the same standard response of the blade element, and such that at least one other basic stress field, associated with this same standard response, presents, with the stress field considered, a dissimilarity of mechanical response lower than another given threshold; during the stage of selecting one or more seeds: • A first seed is selected, for which the mechanical response of the blade element, in response to the stress field corresponding to this first seed, is closest to an average of the mechanical responses of said element subjected to the different basic stress fields associated with the standard response considered, • Next, an additional seed, selected to maximize the dissimilarity of mechanical response compared to that of the seeds already selected, is added to the seed or set of seeds already selected, for which the dissimilarity The mechanical response with said additional seed is the smallest, • this addition operation being executed again until a given total number of seeds is obtained; the seeds are selected so as to be distributed homogeneously throughout all the constraint fields; during the step of adding one of the basic stress fields to the growing subset, the added stress field is, among the basic stress fields that are associated with the same typical response as the seed of the subset, and that are not yet part of the subset, the one that exhibits the smallest mechanical response dissimilarity with respect to the seed of the subset; The method includes: • a determination of baseline training data as defined above, and • a determination of additional training data, in accordance with the method described above, • the basic constraint fields of the basic training data being grouped into groups each associated with one of the said standard responses, on the basis of a dissimilarity of mechanical responses which is the same as the dissimilarity of mechanical response occurring for the formation of pure subsets; the basic training data are determined from high-resolution training data; the high-resolution training data gathers high-resolution stress fields exerted on the turbine blade element and, for each of these fields, an associated label identifying one of said typical responses, each high-resolution stress field giving values of said stress at different nodes of a high-resolution mesh of said blade element; the method includes a step of selecting a limited number of nodes from the high-resolution mesh, the selected nodes forming a relevant node group, the basic constraint fields of the basic training data being obtained by restricting the high-resolution constraint fields to said relevant node group; Determining high-resolution training data involves the following steps: • for each stress field in a set of high-resolution stress fields, said stress field providing values of said stress at the different nodes of a high-resolution mesh: calculation by numerical simulation of a high-resolution mechanical response of said turbine blade element subjected to said stress field, • grouping of the high-resolution mechanical responses thus obtained, in the form of several groups of mechanical responses, each grouping mechanical responses that are similar to each other, based on said dissimilarity of mechanical response, • for each group of mechanical responses, determination of a standard response, representative of the mechanical responses of said group, said standard response being a simplified mechanical behavior model, parameterized by a number of degrees of freedom smaller than the number of degrees of freedom used during said numerical simulation; the simplified mesh corresponding to the relevant node group is such that the relevant nodes are distributed with a substantially homogeneous spatial density over the entire turbine blade element; at least part of said relevant nodes are selected each on the basis of an approximate value of mutual redundancy information between the node in question and any of the other nodes of the high-resolution mesh, said approximate value being determined on the basis of a spatial distance between the two nodes in question; each relevant node is further selected on the basis of a degree of relevance of the value of said stress at the node considered, with respect to the mechanical behavior of said blade element, said degree of relevance being representative of a level of dependence between, on the one hand, the label associated with the high-resolution stress field considered, and, on the other hand, the value of said stress at the node considered for this stress field; The relevant node group initially contains a first node, which is the node of the high-resolution mesh with the greatest degree of relevance, and then an addition step is performed, during which an additional node is added to the relevant node group, the additional node being selected from among the nodes of said mesh that are not yet part of said group, so as to maximize an equal quantity: to the degree of relevance of the node in question, less an average level of redundancy between this node and the nodes already present in the group of relevant nodes, the addition step then being executed again several times.
[0050] The invention also relates to a numerical simulation method of the mechanical response of a turbine blade element subjected to a given stress field comprising: • a preliminary training phase including: • a determination of baseline training data and supplementary training data as explained above, • training a constraint field classifier based on said basic training data and said additional training data, and • a usage phase comprising the following steps: • identification of one of the standard responses for said blade element by the constraint field classifier receiving said constraint field as input, • numerical simulation of the mechanical response of said blade element, subjected to said stress field, in accordance with the simplified mechanical behavior model defined by the standard response previously identified by the classifier.
[0051] The invention also relates to an electronic device comprising at least one processor and one memory, programmed to determine basic training data and additional training data in accordance with the method described above.
[0052] The invention also relates to an electronic device comprising at least one processor and one memory, programmed to execute the numerical simulation method which has just been presented.
[0053] The invention and its various applications will be better understood by reading the following description and examining the accompanying figures. BRIEF DESCRIPTION OF THE FIGURES
[0054] The figures are presented for illustrative purposes only and are in no way limiting of the invention.
[0055] [Fig.1] Fig.1 schematically represents an operation of identification, by a temperature field classifier, of a typical response adapted to a given temperature field, exerted on a turbine blade element.
[0056] [Fig.2] Fig.2 schematically represents a step in determining training data and a training step of a classifier such as that of [Fig.1].
[0057] [Fig. 3] [Fig. 3] schematically represents steps in the determination step of training data from [Fig.2].
[0058] [Fig.4] Fig.4 schematically represents a method for determining training and learning data of a classifier such as that of [Fig.1].
[0059] [Fig. 5] [Fig. 5] Schematically represents, in perspective, an example part mechanical for which the process of [Fig.4] is implemented.
[0060] [Fig.6] Fig.6 schematically represents mutual information values of redundancy between nodes of the mesh of [Fig.5].
[0061] [Fig.7] Fig.7 schematically represents steps for selecting relevant nodes, among the nodes of a complete mesh such as that of [Fig.5],
[0062] [Fig.8] [Fig.8] schematically represents the part of [Fig.5], in perspective, with the set of relevant nodes selected through the steps of [Fig.7].
[0063] [Fig.9] Fig.9 schematically represents the steps of a process data increase.
[0064] [Fig. 10] The [Fig. 10] is a schematic representation, in a reduced two-dimensional space, of two clusters of temperature fields, and of a pure subset identified in one of these clusters.
[0065] [Fig. 11] The [Fig. 11] is a schematic representation, in a reduced two-dimensional space, of different temperature fields each reduced to a point, these temperature fields belonging to different clusters.
[0066] [Fig. 12] The [Fig. 12] schematically represents steps implemented to form a pure subset, such as that of the [Fig. 10].
[0067] [Fig. 13] The [Fig. 13] schematically represents steps implemented to test the purity of such a subset. DETAILED DESCRIPTION
[0068] The embodiment described below corresponds to a case in which the stress exerted on the turbine blade element considered is a temperature exerted on that element (at different points on that element). In other embodiments, the stress in question could be of another type. It could, for example, be a pressure exerted on that blade element by a fluid flowing in the turbine.
[0069] Figure 4 represents a method for determining training data and training a temperature field classifier, comparable to that described in the section relating to the technological background (with reference to Figures 1 to 3) but also incorporating: • an SO step for variable selection, and • an AO step of data augmentation.
[0070] Step 10 of [Fig.4] is the same as that described previously, with reference to Figures 2 and 3. It allows us to determine a first set of training data (TH>i,Mk) bringing together the high-resolution temperature fields TH,i mentioned above, and, associated with each of these temperature fields, the label that identifies the corresponding typical response Mk (i.e. the typical response Mk corresponding to the cluster Ck to which this temperature field belongs).
[0071] The variable selection step S0 is a step in which, for each high-resolution temperature field THi, a corresponding basic temperature field Ti is determined; the basic temperature field Ti provides temperature values at different nodes of a simplified mesh that comprises fewer nodes than the high-resolution mesh. This simplified mesh consists of a group of relevant nodes S, selected from among the different nodes of the high-resolution mesh. This simplified mesh may, for example, comprise a number Q of nodes 100 times, or even 400 times smaller, than the number of nodes QH of the high-resolution mesh.
[0072] The simplified mesh is determined here so that the basic temperature fields T;, significantly simplified compared to the high-resolution temperature fields TH.i, are nevertheless as relevant as possible from the point of view of the thermomechanical response of the turbine blade element 1'; 1 whose behavior we want to simulate.
[0073] Step S0 thus makes it possible to obtain a basic training dataset (T;, Mk), which brings together the basic temperature fields T; in question, each with its label, this label being the same as for the high-resolution temperature field TH>1 which has been "compressed" in the form of the basic temperature field T; considered.
[0074] Next, during the data augmentation step A0, additional training data (T'i, Mk) is determined from the basic training data (T, Mk). The additional training data (T'i, Mk) consists of additional temperature fields T'i, each with a label that identifies one of the typical responses Mb...Mk,...,MP of the turbine blade element.
[0075] Next, in step 20, the temperature field classifier is trained with the basic training data (T;, Mk) and the additional training data (T';, Mk) obtained previously.
[0076] The following section describes in more detail the SO step of variable selection, with reference to Figures 5 to 8. The AO step of data augmentation is then presented, with reference to Figures 9 to 13.
[0077] Step SO: variable selection
[0078] Figure 5 shows, in perspective, the mechanical part 1, which serves here as an example to illustrate the implementation of this method for determining training data. This figure shows the high-resolution mesh, with QH nodes, mentioned above, as well as the positions of 2 and three nodes of this mesh. In this example, the number QH of nodes in the high-resolution mesh is equal to 42445.
[0079] Inset (c) of [Fig. 8] shows, in the form of thick dots, the positions of the different nodes of the relevant node group S obtained for the same For example, using the variable selection technique in question. In this example, the number Q of relevant node group S is equal to 87.
[0080] Here, the relevant nodes are selected: based on mutual redundancy information Ijj between nodes, and based on a degree of relevance y ( j) of each node, with respect to the assignment of such and such a typical response Mk to a given temperature field.
[0081] The mutual redundancy information Ijj • is representative of a level of information redundancy between the temperature value at a first node j of the high-resolution mesh, and the temperature value at a second node j' of this mesh, for the temperature field distribution corresponding to the set of high-resolution temperature fields.
[0082] The relevant nodes are selected so as to have a high degree of relevance, while being little redundant with the nodes already selected, which are already part of the group of relevant nodes S.
[0083] To this end, the relevant node group S is formed as follows. First, a first node, which is the node j of the high-resolution mesh with the highest degree of relevance j ( 'pJ j), is integrated into the relevant node group S. Then, a node addition step is executed several times successively. During this step, an additional node, which is selected, is added to the relevant node group S: • among the nodes of the high-resolution mesh that are not yet part of said group, and
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092] • so as to maximize a quantity equal to the degree of relevance y ( of the node i considered, less an average level of redundancy between this node and the nodes already present in the group of relevant nodes S. More precisely, the additional node, added to the relevant node group S during the addition step, is node i that maximizes the following quantity: ' Card(s) e where i is the index identifying the node in question, in the high-resolution mesh, and where Cârd(S) is the number of nodes already contained in the relevant node group S during the current iteration of the addition step. Here, the iterations of the node addition step cease when the quantity above no longer changes significantly over several successive iterations. For example, if this quantity has not changed by more than 5% (in relative value) over the last 5 iterations, this addition step is no longer repeated. Alternatively, we could also stop repeating this addition step when the quantity in question no longer decreases significantly from one iteration of the step to the next. We could also plan to stop the iterations when the quantity in question falls below a given threshold, or when a predetermined number of iterations is reached. We could also plan to stop the iterations when the relevance information j of the node to be added falls below a given threshold. This criterion is interesting because the quantity to be maximized at each iteration (the quantity above) can take on a high value for a node with low relevance information Ip(T1) (and therefore low relevance) but also with low redundancy with the already selected nodes. As already mentioned, the mutual redundancy information Ijj • is representative of the The level of information redundancy between the temperature value TJ at a first node j of the high-resolution mesh, and the temperature value Tj- at a second node j' of this mesh. More precisely, this mutual redundancy information, which reflects a greater or lesser degree of interdependence between the temperatures Tj and Tj, is defined by the following relationship: TJ ) =SPjj.(f, tJ')logl dt' dt1 Or : - the integration domain is R2,
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103] - p is the probability density function for the temperature field distribution given by the set of high-resolution temperature fields TH> b ..., THN, to have a temperature tj at node j, - p) is 'the probability density, for this same distribution, of having an ay' temperature tj at node j', and where - p') is a (joint) probability density function, for this same distribution, to have a temperature tj at node j and a temperature tj at node j'. As for the degree of relevance j (yf j of node j, with respect to the assignment of such and such a response type Mk to a given temperature field, it is representative of a level of dependence between: • on the one hand, the label number (i.e., the index k identifying the appropriate standard response), associated with a given high-resolution temperature field, and, • on the other hand, the temperature value at node j, given by this temperature field. The degree of relevance ] [ pJ'j is a particular form of mutual redundancy information, in the case where a first variable is the temperature Tj at node j, while the second variable, Y, is discrete and represents the label number associated with the temperature field: j ( jd) = f ( pi y) with Or : - the integration domain is R, - p ( is the probability density function, for the temperature field distribution) given by the set of high-resolution temperature fields TH> b ..., THN, to have a temperature tj at node j, -p#) is the probability, for a standard response to be Mk, ( is a probability density function for the distribution of fields of temperature given by the set of high-resolution temperature fields TH> i, ..., Thn, to have a temperature tj at node j knowing that the standard response knowing that the standard response is Mk,
[0104] P(Hk) is the probability density for the temperature field distribution given by the set of high-resolution temperature fields TH>i, Thn, of having a temperature tj at node j and for the standard response to be Mk, simultaneously.
[0105] The degree of relevance [ ( ■> ct the mutual information of redundancy Ijj • can be determined, on the basis of the above expressions, after a prior estimation of the different probability functions involved in these expressions.
[0106] These probability functions can for example be estimated, for the (discrete) distribution of temperature fields and labels constituted by the set of high resolution temperature fields TH, i, ..., THjn, with their labels, by "binning" techniques (i.e. by sampling techniques).
[0107] By way of example, the following article describes such a "binning" technique, which can be used to calculate such mutual redundancy information from discrete data: "Estimating mutual information", A. Kraskov, H. Stôgbauer, and P. Grassberger, Phys. Rev. E 69, 066138 - 23 June 2004; Erratum: Phys. Rev. E 83, 019903 (2011).
[0108] In any case, a determination of the mutual redundancy information Ijj • between two nodes j and j', by such a statistical calculation based on the temperature values TjHj, ..., TjH>N, and Tj Hj, ..., Tj HN at the level of these two nodes, is computationally intensive.
[0109] And the number of terms Ijj • to be calculated is very large since it is the number of pairs grouping any two nodes j, j' of the high-resolution mesh. In the numerical example corresponding to [Fig.5], the number of pairs is on the order of 109, for example.
[0110] Also, to reduce the computation time required to evaluate the mutual redundancy information Ijj • between nodes, rather than determining it by direct statistical calculation (which would be very long), an approximate value J( dj j,) of this quantity is determined here from the spatial distance, djj •, between the two nodes j and j' considered.
[0111] In other words, we consider here that the mutual redundancy information Ijj • can be expressed, up to a good approximation, in the form of a function of the . This distance djj • is equal to the distance Geometric II FF., It separates the two nodes j, j' in question. This distance It JJ 1'2 is quick to calculate, from a numerical point of view. Evaluating the mutual redundancy information Ijj in this way allows us to form the relevant node group S of distance djj • in question ■ I > I( dji J rj \ Jfj a much faster way than evaluating all the terms Ijj • by direct statistical calculation.
[0112] The one-variable function I, which is used to calculate the approximate values J (dj j J from the distances djj -, is determined beforehand, during a calibration step SI 1 (figure 7), by determining the mutual redundancy information Ijj • in an "exact" way, that is to say by direct statistical calculation, for some nodes of the high-resolution mesh, in order to calibrate the curve, or, in other words, the function ï.
[0113] The SI 1 calibration step comprises the following steps: • S12: for each pair of nodes (j, j') of a subset grouping a portion of the nodes of the high-resolution mesh: • determination of a reference value yJ y / J (the "exact" value) of the mutual redundancy information Ijj • between the two nodes j and j' considered, by statistical calculation based on the temperature values TjH>i, ..., TjH.N, Tj H,i, ..., Tj h,n , at the level of these two nodes, as given by the different high-resolution temperature fields TH> b ..., THN, and • calculation of the spatial distance djj separating the two nodes j and j' considered, then • S13: adjustment of the curve ï so that it is representative of a link between the reference values j ( y J y / ' j thus obtained, and the corresponding spatial distances djj.
[0114] The subset of nodes, for which the mutual redundancy information is calculated in an "exact" manner, may include a number of nodes, for example, between one thousandth and one twentieth of the total number QH of nodes in the high-resolution mesh.
[0115] Curve I is a curve parameterized by one or more parameters which are adjusted (for example by a least squares method) so that this curve is as representative as possible of the reference data
[0116] Figure 6 shows reference values j [ <pi J de l’information mutuelle de redondance Ijj • entre nœuds, obtenues pour la pièce 1 de la figure 5. Sur cette figure, chaque point représente un couple (djj - / ( Cette figure montre also the curve I obtained by fitting these reference data.
[0117] In this example, approximately 750 reference values j (fj j'j' j) were determined to calibrate the curve î. In this case, the curve î represents the following single-variable function:
[0118] i(r)=iæ+ yr(r1-r)a\H(r1-r) + y2.(r2-r)a2.H(r2-r)
[0119] where H is the Heaviside stair-step function, and where Ioo, y-Tp Clp yr2, Ct2 are the parameters adjusted during step S13.
[0120] As can be seen from this example, the mutual redundancy information Ijj • is well correlated with the distance djj • , and can indeed be expressed, to a good approximation, in the form of a function ï of this distance.
[0121] It may be noted that this result is far from obvious at first glance, if only because the mesh of the part is a three-dimensional mesh, which is generally neither homogeneous nor isotropic. Surprisingly, however, it turns out that the mutual redundancy information Ijj • between two nodes j and j' of this mesh can be suitably expressed as a function depending only on the distance djj • between these nodes (whereas the vector connecting these two nodes can be oriented in many different ways, for example).
[0122] This property, exploited here in a particularly interesting way, can be explained as follows. The temperature fields used for numerical simulations of the mechanical behavior of turbine blades are generally at least partially random fields with Gaussian statistics (this is notably the case for the high-resolution temperature fields used here). The temperatures at two nodes j and j', as given by this set of fields, can then be seen as Gaussian random vectors. And it can be mathematically shown that the mutual redundancy information j(x\ X2) between two such vectors X1 and X2 is expressed as:
[0323] j(x\ X2)= 4 ln(l-p2)
[0124] where p is the correlation between X1 and X2.
[0125] Moreover, these random temperature fields generally have isotropic correlation functions. The correlation p then depends only on the distance djj • between the nodes considered, which finally explains why the mutual redundancy information Ijj • can be expressed as a function of the single distance djj -.
[0126] In any case, at the end of the SI 1 calibration step, the curve î is completely determined, to then be used for the selection of relevant nodes.
[0127] The relevance degree f(pQ) is calculated during step S14 for all nodes of the high-resolution mesh. It is determined by statistical calculation. direct. The computation time required to carry out this step remains reasonable, however, because, unlike the case of mutual information, it involves calculating a value per node (and therefore QH different values), and not a value per pair of nodes (i.e. Qh.(Qh- l) / 2 different values).
[0128] Then, in a step S15, the relevant node group S is determined, as explained above, by selecting relevant nodes based on the degree of relevance and mutual redundancy information. As the relevant node group S grows progressively, the additional node added to the nodes already present in this group is node i that maximizes the following quantity: 101291
[0130] As already indicated, inset (c) of [Fig. 8] shows the different nodes of the relevant node group S obtained for the high-resolution mesh example of [Fig. 5]. In this example, the node addition iterations were stopped once the above quantity ceased to vary significantly from one iteration to the next. 87 relevant nodes were thus selected (which shows, incidentally, that significant compression of the initial high-resolution mesh was possible). 303 seconds were required to identify this group of 87 relevant nodes S (i.e., to perform 86 successive executions of the addition step). By comparison, when the mutual redundancy information Ijj• is evaluated by direct statistical calculation, it takes, on the same computer, 6469 seconds to execute only the first 6 iterations of this selection algorithm (and thus select 7 relevant nodes).This illustrates the very significant acceleration made possible by the approximate calculation technique for mutual information presented here.
[0131] Inset (a) of [Fig. 8] shows a group _S' of 87 nodes, selected from the high-resolution mesh using a different selection technique. In this case, the nodes in this group were selected based on a purely geometric criterion, so as to obtain, for the selected nodes, a substantially homogeneous node density over the entire part (for example, such that the volume of a zone specific to each selected node, centered on that node and containing no other relevant node, varies little from one selected node to another). This selection is performed without taking into account the mechanical responses of the part.
[0132] Inset (b) of [Fig. 8] shows a group _S” of 87 nodes, selected from the high-resolution mesh using yet another selection technique, based on a univariate filter. As can be seen in this figure, the nodes selected in this way are concentrated mainly at a central constriction in part 1, and on the surface of this part (an area that is a priori quite critical from a thermomechanical point of view). On the other hand, we observe that virtually no nodes are selected anywhere other than in this area, and that many nodes are located (probably unnecessarily) in the same place.
[0133] The selection technique presented above, based on a degree of relevance corrected for redundancy, leads to a distribution of nodes (inset (c)), which, compared to the homogeneous distribution of inset (a), is denser at the central narrowing, while covering the whole of this piece (unlike selection by univariate filter).
[0134] The overall relevance degree D of the group of relevant nodes S, S' and S”, and its overall redundancy level R are given below, for these three selection techniques: • (a): purely geometric distribution; D=0.009 and R=0.1072; • (b): selection by univariate filter; D=0.0671 and R=0.8129; • (c): maximizing relevance while minimizing redundancy; D=0.046 and R=0.111.
[0135] A purely geometric distribution therefore leads to a low redundancy R, but also to a low relevance D.
[0136] Selection by univariate filter leads to high relevance D, but to high redundancy R.
[0137] As for selection by maximizing relevance and minimizing redundancy, it leads to both high relevance D and low redundancy R.
[0138] The overall relevance degree D of a group of relevant nodes given above is defined by the following formula:
[0139] 0 =
[0140] The overall redundancy level R of a group of relevant nodes is defined by the following formula: 101411
[0142] For the method of determining training data described here, the selection technique, used during the step of selecting variables S0, is technique (c), maximizing relevance while minimizing redundancy.
[0143] Alternatively, another selection technique, for example by univariate filter or by purely geometric distribution, could however be used during step S0.
[0144] In any case, at the end of the variable selection step, we have a set of basic temperature fields Tb.. .,T;,.. ,,TN, with their labels, less voluminous than the initial high-resolution temperature fields TH b ... ,TH.i,... ,Th.n.
[0145] During the next AO step, additional training data are determined, by data augmentation, from these basic training data.
[0146] And ape AO: data augmentation
[0147] The AO data augmentation step comprises the following steps ([Fig.9]): • Step A1: form different subsets pure Sb and • Step A2: within each pure subset, determine data additional learning through convex linear combinations.
[0148] At step Al, each pure subset Si is formed in such a way: • to group together basic temperature fields Tj that are associated with the same typical response Mk of the turbine blade element (i.e., that belong to the same cluster Ck), • and so as to have a convex envelope Ei which does not contain any basic temperature field T; which is associated with another response-type Mk -, kVk, of the turbine blade element (i.e. no basic temperature field T; foreign, belonging to a cluster other than the Ck cluster).
[0149] Then, during step A2, for each subset Sb, different additional training data are determined, each additional training data comprising: • one of the additional temperature fields T'i in question, which is obtained by linear combination of the basic temperature fields T; of the pure subset Si considered, assigned positive weighting coefficients and normalized to 1, and • a label (for example, the index k identifying the standard response Mk), which is the same as for the basic temperature fields T; which were used to obtain this additional temperature field T'b
[0150] Step A1, for the formation of pure subsets Sb 1=1..M, is now described in more detail. This step comprises the following steps: • Step A10: among the base temperature fields T, belonging to the same cluster Ck, selection of "seeds" SDb 1=1.. .M, for the subsequent growth of pure subsets Sb • Step Al 1: for each seed SDb, formation of a subset Sb initially containing said seed, and which is grown until its purity is lost.
[0151] The A10 seed selection step is presented first, and the Garlic growth step, which includes a specific purity test, is described next.
[0152] Step A10: Seed selection
[0153] Each SDi seed consists of one of the basic temperature fields Ti belonging to the considered cluster Ck, judiciously selected.
[0154] Here, the SDi seeds are chosen, within the considered cluster Ck, so as to be somewhat distant from the cluster boundaries. More precisely, they are chosen so as to be a minimum distance from the basic temperature fields belonging to the other clusters C k • k' k. This distance is a distance in the sense of the distance described above (in the part relating to the general technological background), that is to say, a distance which reflects a dissimilarity of the mechanical responses, obtained respectively for the two basic temperature fields 1), Tj considered (dissimilarity of mechanical behaviors).
[0155] Thus, in each cluster Ck, the SDi seeds are selected so as to exhibit, with respect to the different basic temperature fields Tr of the other clusters Ck -, k'^k, a dissimilarity ôy • of mechanical response greater than a given threshold ôi.
[0156] In each cluster, it is advantageous to choose seeds far from the base temperature fields of other clusters Ck • , because they are conducive to the subsequent growth of a pure subset Sp Indeed, thanks to this distance, this subset will be able to grow a minimum in size before a base temperature field, belonging to another cluster Ck •, finds itself in its convex hull Ei.
[0157] More generally, this distance limits the risk that the convex hull Ei of the subset Si will encompass, in the space of temperature fields, a part of another Cluster Ck•, whether or not that part contains a basic temperature field. In other words, for the example in [Fig. 10], choosing seeds in the cluster Ci that are far from the crosses representing the basic temperature fields of the cluster C2 prevents a subset, originating from one of these seeds, from encompassing a part of the domain corresponding to the cluster C2, a domain represented in white in the figure, whether or not that part contains a cross.
[0158] Furthermore, within each cluster Ck, the seeds are chosen so as not to be too isolated from the other basic temperature fields of that same cluster, so that the seed is conducive to the subsequent growth of a subset. More precisely, each seed is chosen so that at least one other basic temperature field of the same cluster Ck is located at a "distance" ôÿ from the seed that is less than a given threshold ô2 (in other words, which exhibits, with the seed in question, a dissimilarity ôÿ of mechanical response less than the threshold ô2 in question).
[0159] To obtain seeds that are not too isolated from the other base temperature fields in their cluster Ck, and that do not have close neighbors belonging to a foreign cluster Ck, we begin by selecting, in the form of a list Fk, all the base temperature fields that satisfy these two criteria in the cluster Ck under consideration. This list is obtained by removing (deleting) from the complete list grouping all the base temperature fields of the cluster Ck under consideration: • all basic temperature fields 1), having a neighbor Ti' which belongs to another cluster and which, with respect to 1), is located at a distance by • less than the threshold ôb and • all base temperature fields T; which have no neighbor Tr belonging to the same cluster and located at a distance by from T, less than the threshold b2
[0160] The thresholds ôi and b2 are defined here as being equal to Ei.bref and e2.ôref respectively, where ôref is a distance representing a dimension of the cluster Ck under consideration (the values of the thresholds used may therefore vary from one cluster to another). The coefficients Ei and e2 are less than 1 (and positive). For example, Ei can be between 0.01 and 0.05, and e2 can be between 0.1 and 0.3 (this means that we also select seeds that are relatively "isolated". This is useful because the data augmentation also aims to generate data in areas that lack it, i.e., those that are poorly represented by the initial data).
[0161] ôref is, for example, determined to be the distance between: • a base temperature field of the cluster Ck, considered to be representative of that cluster, and • the base temperature field, foreign to the Ck cluster, and which is closest to the representative of Ck.
[0162] The representative of the Ck cluster is, for example, the basic temperature field whose mechanical response is closest to an average of the mechanical responses obtained for the different basic temperature fields of this cluster.
[0163] Once the list Fk of temperature fields in the cluster Ck that are neither too isolated nor too close to the boundaries has been obtained, a given number of seeds are selected from this list. This number is, for example, previously set by a user.
[0164] The seeds can be selected, from the Fk list, so as to be distributed evenly in the cluster.
[0165] Here, the seeds can be selected as follows, from the Fk list. The first SDi seed selected is the base temperature field, which is representative of the Ck cluster under consideration. The other seeds are selected, one after the other. The other method ensures that each newly selected seed is comprised of the base temperature field in the Fk list that is furthest from the nearest seed. The second seed, SD2, for example, is comprised of the base temperature field in the Fk list that is furthest from the first seed, SDi. And the third seed, SD3, is comprised of the base temperature field in the Fk list that is furthest from either the first seed, SDi, or the second seed, SD2, depending on whether the nearest seed to SD3 is SDi or SD2.
[0166] Thus, if Lk denotes the current list of selected seeds, the next seed added to this list is the base temperature field Tj whose index j is given by the following formula:
[0167] j=arg(max [min<5,hl) \aEFk\Lk[bELk J
[0168] where ô is the distance between temperature fields mentioned above (also called mechanical response dissimilarity).
[0169] Selecting the seeds in this way allows for homogeneous coverage of the entire Ck cluster, and therefore, subsequently, the creation of pure subsets that are also homogeneously distributed within the cluster. This is useful because pure subsets distributed inhomogeneously would lead, after the creation of additional training data within these subsets, to an overrepresentation of certain areas of the cluster, which is not optimal for classifier training.
[0170] Garlic step: growth of pure subsets.
[0171] Once the SDi seeds, 1=1..M, have been selected, for each SDi seed, a corresponding subset Si is grown from the seed in question.
[0172] Initially, the subset Si contains only this seed SDb. It is then grown by performing the following steps ([Fig. 12]): • A10: adding one of the basic temperature fields Ti to the subset Si, the added temperature field having the same standard response Mk as the other temperature fields of the subset Si (in other words, the added temperature field belongs to the same cluster Ck as the other elements of SJ, then • Al 11: purity test, for the subset Si thus augmented.
[0173] If step Al 11 shows that the subset Si, thus augmented by another base temperature field, is still pure, step Al 10 is executed again, then step Al 11, and so on until step Al 11 shows that the subset Si is no longer pure. The iterations of steps Al 10 and Al 11 are then stopped, and the last base temperature field to have been added to the subset is removed from this subset Si, which makes it pure again.
[0174] The growth of the subset Si is therefore carried out step by step. More precisely, during the addition step Al 10, the base temperature field T; which is closest to the seed SDi (in the sense of distances ôÿ), is added to the subset Si, among the base temperature fields of the cluster Ck which have not yet been selected (i.e., which are not already part of the subset SJ.
[0175] It should be noted that the subsets Si are thus grown step by step, based on a particular metric ôÿ which represents the dissimilarity between the mechanical responses associated with the temperature fields, and not based on a usual (e.g., Euclidean) distance between the temperature fields themselves. This allows each subset to grow with a good probability of remaining pure (since the purity criterion in question concerns the mechanical responses associated with the temperature fields). In other words, growing the subsets Si based on the same metric ôÿ as that initially used to identify the different clusters Ch C2, ... CP (i.e., the different classes) of temperature fields optimizes the probability that the subset remains pure during its growth. This therefore makes it possible to obtain larger pure subsets.This precaution is welcome in practice, because even when growing the pure subsets in this way, their maximum size generally remains small (they typically contain only a few temperature fields, on the order of 5 to 10, on average, for the examples presented here).
[0176] Similarly, the fact of pre-selecting the seeds (at step A10) on the basis of this same metric ôÿ, initially used to identify the different clusters Ci, C2, ... CP, makes these seeds good candidates for the subsequent growth of pure subsets.
[0177] Purity Test
[0178] During the purity test step Al 11, it is verified, for the subset Si to be tested, that no basic temperature field Tr, associated with a standard response Mk • different from the standard response Mk common to the basic temperature fields T, of the subset Si considered, is located inside the convex envelope Ei of this subset Si.
[0179] In other words, it is verified that no basic temperature field, belonging to a cluster Ck other than the cluster Ck of the subset to be tested, is located inside the convex hull Ei of the subset Si. When this condition is satisfied, the subset in question is considered pure.
[0180] This test amounts to verifying, for each basic temperature field Tr belonging to another cluster Ck, that this temperature field Tr cannot be expressed as a linear combination of the basic temperature fields Ti belonging to the subset Si in question, this linear combination having positive coefficients and normalized to 1.
[0181] The criterion to be tested is therefore relatively computationally intensive, even if the number of degrees of freedom of the basic temperature fields T; is relatively reduced, compared to the initial high-resolution temperature fields TH i (thanks to the prior step of variable selection S0).
[0182] Also, to reduce the computation time required to perform the purity test Al 11, the purity of each subset Si is tested by projection, i.e. on the basis of temperature subfields Tj, obtained by projecting the basic temperature fields Ti onto a subspace with a number of degrees of freedom less than the number Q of nodes of the mesh on which the basic temperature fields T are defined.
[0183] As indicated in the description of the variable selection step S0, each base temperature field Ti groups together a set of Q temperature values T;1,.. ..,TiQ , which are the temperature values of that field at the different nodes of the mesh in question (which, in this case, is an already simplified mesh, compared to the initial high-resolution mesh which has QH nodes).
[0184] The subspace onto which the basic temperature fields Tj are projected is a QR-dimensional space, that is, one with a number of scalar degrees of freedom QR less than Q (in practice, it may be a subspace of very small dimensions, for example, only 5 dimensions instead of 87 or 42445 dimensions). This subspace corresponds here to the set of temperature fields that give temperature values at only a number QR of mesh nodes, instead of giving these values at all the mesh nodes on which the basic temperature fields Tj are defined.
[0185] Here, each temperature subfield jL thus gives temperature values at different nodes of a restricted node group, comprising only QR nodes (for example, only 5 nodes) instead of Q nodes. These nodes are chosen from among the nodes of the (reduced) mesh on which the basic temperature fields T are defined; that is, from among the Q nodes of the relevant node group S determined in step S0. In this case, they are chosen randomly from among the Q nodes of the relevant node group S.
[0186] Each temperature subfield Tj is then obtained from the base temperature field T; corresponding, while preserving, among the temperature values T / ,.. ..,T;q given by Ti, a restricted number of these temperature values, which are those at the level of the nodes of the restricted group in question.
[0187] It should be noted, however, that other ways of projecting the temperature fields T; are conceivable. Any projection onto a subspace of RQ is conceivable. This subspace may, for example, correspond to the QR first components of the canonical basis of RQ, or to the QR first left singular vectors in the singular value decomposition (SVD) of the collection of temperature vectors T / ,.. ..,T;Q at the nodes of the reduced mesh..
[0188] In any case, carrying out the Al 11 purity test mentioned above, not on the basis of the basic temperature fields T;, but on the basis of the less voluminous subfields Tp, makes it possible to accelerate this purity test step.
[0189] Moreover, surprisingly, when the result of this purity test by projection is positive, this validates the purity of the subset Si to be tested as reliably (as far as purity validation is concerned) as if the test had been carried out on the basis of the complete basic temperature fields.
[0190] Indeed, a mathematical demonstration shows that: • if the convex hull Ëj of the set grouping the temperature subfields that are the projections of the basic temperature fields T; belonging to a subset Si, contains no temperature subfield that is associated with a basic temperature field Tr of a cluster other than that of the subset Si, • then, the convex envelope Ei of the subset Si does not contain any base temperature fields Tr of a cluster other than that of the subset Si.
[0191] That is to say, in summary: if the projection of the subset Si is pure, then the subset Si itself is also pure.
[0192] The demonstration of this property is not included here for the sake of brevity, and because it is not of direct interest for the implementation of the invention.
[0193] It should be noted, however, that the converse of this proposition is generally not true. Indeed, if the projection of the subset Si is not pure, one cannot generally deduce anything from it concerning the purity or impurity of the subset Si itself.
[0194] Thus, the projection of the subset Si may very well not be pure, even though the subset itself is pure. Indeed, it is clear that a projection, which constitutes a kind of compression along certain directions, can very well bring into a subset elements belonging to another cluster, which, before the projection, were outside that subset. In other words, a projection can very well cause a subset to lose its purity.
[0195] To take into account these two aspects (namely: the purity of the projected implies the purity of the subset Si itself; and the non-purity of the projected implies uncertainty as to the purity of the subset Si), the purity test step Al 11 is carried out as follows ([Fig. 13]): • we execute a test step by projection Al 12; • when the result of the Al 12 projection test is positive (i.e., when the projected subset Si is pure), we determine that the subset Si is pure; • On the other hand, when the result of the Al 12 projection test is negative, we perform this projection test, but by carrying out the projection on the basis of another subspace (of another restricted group of nodes, selected randomly).
[0196] As long as the result of the projection test Al 12 is negative, and as long as the number it of iterations of this test is less than a limit number of iterations, niim, the step Al 12 is executed again.
[0197] If the result of the Al 12 projection test eventually becomes positive, during these iterations, it is determined that the subset Si is pure.
[0198] Conversely, if after niim iterations of the projection test (each time projecting onto a different subset), the test still does not yield a positive result, it is determined that the subset Si to be tested is not pure. In reality, it is unknown whether it is pure or not. However, since it could not be ascertained that it is pure, this subset is eliminated as a precaution, and no additional temperature field will be determined within it.
[0199] In practice, the Al 12 projection test can be carried out as explained now.
[0200] During the Al 12 projection test, it is determined, for the subset Si whose purity is to be tested, whether the convex envelope Ëj of the set grouping the temperature subfields Tp which are the projections of the basic temperature fields T; of the subset Si, does not contain any temperature subfield which is associated with a foreign basic temperature field Tr, of a different cluster Ck than the cluster Ck to which the subset Si to be tested belongs.
[0201] This amounts to testing that no temperature subfield can be written as a linear combination of the temperature subfields Tj of the subset Si, assigned positive coefficients and normalized to 1. To test this criterion, we proceed as follows.
[0202] For each temperature subfield T corresponding to one of the basic temperature fields T of the subset Si, a vector jt is considered, having QR+1 components. The first QR components of the vector jq are the QR temperature values given by the temperature subfield. The last component of this vector is equal to 1.
[0203] We then form a matrix A, of dimension (QR+l)xn, by grouping in matrix form the n column vectors jt; corresponding to the n temperature subfields Ti that comprise the subset Si whose purity we want to test.
[0204] For each temperature subfield belonging to a different cluster Ck- than that of the subset Si to be tested, a vector ny is constructed in the same way
[0205] To test whether the temperature subfield T is indeed different from any linear combination of the temperature subfields Tj of the subset Si to be tested, assigned positive coefficients and normalized to 1, we then test whether the condition below is satisfied:
[0206] min II Aw-nr|| > &||n>ll weR"
[0207] where w is a vector grouping n positive coefficients and where e is a fixed tolerance level, for example between 0.5% and 5%.
[0208] From a numerical point of view, to test this condition, a least-squares minimization (with positive coefficients) is performed on the left-hand side. The coefficients grouped in the vector w are not normalized to one (removing this constraint simplifies the minimization procedure in question). However, the component equal to 1, added to the end of each vector ir, allows this test to be performed as if testing a linear combination with positive coefficients normalized to 1.
[0209] Numerical example
[0210] Results obtained using this method of determining training data are presented briefly below as an example of implementation.
[0211] This example corresponds to the part shown in [Fig. 5], for which the high-resolution mesh comprises QH = 42,445 nodes. Furthermore, N = 600 high-resolution temperature fields TH15..., TH N are initially produced in step 10. And the number P of clusters (i.e., the number P of Mk response types) is equal to 4 here.
[0212] As already indicated, in this example, the variable selection step S0 allows, in the high-resolution mesh, the selection of Q=87 relevant nodes, which allows the high-resolution temperature fields to be simplified into the basic temperature fields Tb..., TN, still numbering 600.
[0213] Next, during the AO data augmentation step, approximately 60 pure subsets were identified (on average) in each Ci,...CP cluster of temperature fields. Each pure subset thus identified comprises on average 5 elements (5 basic temperature fields). During the purity test, by projection (step Al 12), the subspaces onto which the temperature fields are projected are 5-dimensional subspaces (in other words, 5 nodes, instead of 87, for the purity test by projection).
[0214] In this example, 5400 additional temperature fields T'i, each with its label, were then obtained by convex linear combination within these pure subsets.
[0215] The set of training data obtained (basic + additional) then includes 6000 labeled temperature fields, used for training a temperature classifier.
[0216] The classification accuracy obtained after training was then tested for several types of classifiers: • with and without data augmentation (with / without DA), and • with a selection of variables: • either by the selection technique described above, by maximizing relevance while minimizing redundancy (FS), • either by principal component analysis (PCA).
[0217] These results are summarized in the table below. Classifier Type Used Dimension Reduction Accuracy, With DA Accuracy, Without DA Multilayer Perceptron FS 87.0% 81.0% PCA 86.5% 81.5% Quadratic Discriminant Analysis FS 77.5% 70.5% PCA 76.0% 70.0% Gaussian Naive Bayes FS 39.5% 34.5% PCA 38.5% 31.5%
[0218] These results show that additional training data, as determined here, do indeed improve classification accuracy, and are therefore physically relevant.
[0219] These results also show that the variable selection (FS) method presented above is reliable (in addition to being fast to execute), in that it allows effectively extracting the most relevant information from high-resolution temperature fields.
[0220] Electronic devices
[0221] The present technology also relates to an electronic device comprising at least one processor and one memory, programmed to execute the variable selection method and / or the data augmentation method described above. More generally, this electronic device can be programmed to execute the entire method for determining training data and training the classifier of [Fig. 4].
[0222] The device in question may be implemented in the form of a computer or a computing unit whose components are located in the same place, for example in the same room or in the same computer case. It may also be implemented in the form of a decentralized computing system that uses, via a computer network, remote, distributed computing resources (for example, remote computing resources provided by a "cloud").
[0223] The present technology also relates to an electronic device, comprising at least one processor and one memory, including a classifier, i.e., programmed to perform a classification function as described above, and in which the classifier has been trained on the basis of the basic and / or supplementary training data presented above. This classifier is specific (and, as seen above, particularly accurate) due to the specific training it has undergone. From a practical point of view, the specific characteristics of this classifier are reflected in the particular values of the coefficients (synaptic coefficients, for example) that parameterize the classifier, as obtained after training on the basis of the basic and / or supplementary training data.
[0224] The present technology also relates to an electronic device, comprising at least a processor and a memory, programmed to implement a numerical simulation method of the mechanical response of a turbine blade element subjected to a given temperature field TR, the simulation method comprising: • a preliminary training phase including: • a determination of baseline training data and supplementary training data, as explained above, • training a temperature field classifier based on said basic training data and said additional training data, and • a usage phase comprising the following steps: identification of one of the standard responses Mb ..MP of said blade element by the temperature field classifier receiving as input said temperature field TR, and numerical simulation of the mechanical response of said blade element, subjected to said temperature field TR, in accordance with the simplified mechanical behavior model defined by the standard response previously identified by the classifier.
Claims
37 Claims
1. A computer-implemented method for characterizing the aging of a turbine blade of a turbomachine based on a numerical simulation of the mechanical response of an element of the turbine blade subjected to a given stress field (TR) comprising: - a preliminary training phase comprising the determination of basic training data and additional training data for training a stress field classifier (2'), the additional training data being determined from basic training data, the basic training data comprising basic stress fields (Tl, Ti, TN) exerted on a turbine blade element (1'; 1) and, for each of these fields, a label (k) which identifies a typical response (Mk), representative of an expected mechanical response for this turbine blade element (1'; 1) in response to said stress field (Ti), among different typical responses (Ml, Mk, MP) of said turbine blade element, the determination of said basic training data and said additional training data comprising the following steps: • (Al) form at least one pure subset (Si), which groups basic stress fields (Tr) associated with the same standard response (Mk) of the turbine blade element, and which has a convex envelope (Ei) containing no basic stress field (Tr) associated with another standard response (Mk) of said element, • (A2) for at least one of said subsets pure (Si), determination of said additional training data, each additional training data comprising an additional constraint field (T'i), obtained by linear combination of basic constraint fields (Ti) of the pure subset (SJ considered, assigned positive weighting coefficients and normalized to one, the additional constraint field (T'i) being accompanied
2.
3. of a label (k) which is the same as for the basic constraint fields (Tj) of said combination, and - a usage phase comprising the following steps: • identification of one of the standard responses (Mb Mk, MP) of said blade element by the stress field classifier receiving said stress field (TR) as input, • numerical simulation of the mechanical response of said blade element, subjected to said stress field (TR), in accordance with the simplified mechanical behavior model defined by the standard response previously identified by the classifier. Method according to the preceding claim, wherein said stress exerted on the turbine blade element (1'; 1) is a temperature exerted on this element, the basic stress fields (Ti, T;, TN) and the additional stress fields (T'i) being respectively basic temperature fields and additional temperature fields. Method according to any one of the preceding claims, wherein the step (Al) of forming one or more pure subsets (Si, Si, SM) comprises, for each of these subsets, a purity test step (Al 11) of the subset (Si) under consideration, during which it is verified that no basic stress field (Tr) associated with a standard response (Mk) different from the standard response (Mk) common to the basic stress fields (Ti) of the subset (SJ under consideration, is located inside the convex hull (Ei) of said subset (Si).
4. Method according to the preceding claim, wherein: - each basic constraint field (T;) groups a set of values of said constraint (T;1, Tij,T;Q) into different nodes of a given mesh comprising a total number of nodes Q, - for each basic constraint field (Tj), an associated constraint subfield (T^) is obtained by projecting the basic constraint field (T;) onto a subspace with a number of degrees of freedom less than Q, and - during the purity test step (Al 11): • a projection test (Al 12) is performed, during which it is determined whether, for the stress subfields (Tp) associated with the basic stress fields (Tj) of the subset to be tested, a convex hull (Kp) of the set grouping these stress subfields (Tp) contains no stress subfield (T^) associated with one of the basic stress fields (T; ) for which the standard response (Mk -) is different from the standard response (Mk) common to the basic stress fields (T;) of the subset (SJ) to be tested, • and, when the result of this projection test (Al 12) is positive, it is determined that the subset (SJ of stress fields (1)) considered is pure.
5. Method according to the preceding claim, wherein, during the purity test step (Al 11), - as long as the result of said projection test (Al 12) is negative, this projection test (Al 12) is executed again, the subspace onto which said basic constraint fields are projected being different from one execution of the projection test to another, - as soon as the result of the projection test is positive, we determine that the subset (Si) of stress fields to be tested is pure, - and, when after a given number of successive executions of the projection test (SJ, the projection test (Al 12) has still not given a positive result, it is determined that the subset (SJ to be tested is not pure.
6. A method according to any one of claims 3 to 5, wherein the step (A1) of forming said at least one pure subset (Si) comprises the following steps: - (A10) from among the basic stress fields (Tj) that are associated with the same standard response (Mk) of the blade element, selection of one or more basic stress fields, each constituting a seed (SDi) for the subsequent growth of the pure subset(s) (SJ), - (A11) for each of said seeds (SDi), formation of a subset (Si), initially containing said seed, and which is grown: • by performing an addition step (A10), during which one of the basic stress fields (Tj) is added to the subset (SJ) under consideration, the added stress field being associated with the same standard response (Mk) as the other stress fields of said subset, and then • by performing said purity test step (Al 11),for the subset (SJ thus augmented, • the addition step (Al 10) and the purity test step (Al 11) being executed again, as long as the subset (Si) remains pure.,
7. Method according to the preceding claim, wherein, during the selection step (A10) of one or more seeds (SDi), the seeds are selected from among the basic stress fields (Tj) which are associated with the same standard response (Mk) of the blade element, and which exhibit, with respect to the different basic stress fields (Tj) associated with the other standard responses (Mk), a dissimilarity (ôy-) of mechanical response greater than a given threshold (ôi).
8. A method according to any one of claims 6 or 7, wherein, during the selection step (A10) of one or more seeds (SDi), each seed is selected from among the basic stress fields (T;) which are associated with the same standard response (Mk) of the blade element, and such that at least one other basic stress field (Tr), associated with this same standard response (Mk), exhibits, with the stress field (Tj) considered, a dissimilarity (ôy ••) of mechanical response lower than another given threshold (ô2).
9. Method according to any one of claims 6 to 8, wherein, during the step of adding (Al 10) one of the basic stress fields (Tj) to the growing subset (Si), the added stress field is, among the basic stress fields (Tj) that are associated with the same typical response (Mk) as the seed (SDi) of the subset, and that are not yet part of the subset (Si), the one that exhibits the smallest mechanical response dissimilarity (δ) with respect to the seed (SDi) of the subset.
10. Method for determining training data for training a constraint field classifier, the method comprising: - a determination of basic training data as defined in claim 1, and - a determination of additional training data, in accordance with the method defined by any one of claims 7 to 9, - the basic constraint fields (Tb T;, TN) of the basic training data being grouped into groups (Ck) each associated with one of the said standard responses (Mk), on the basis of a dissimilarity (ô) of mechanical responses which is the same as the dissimilarity (ô) of mechanical response involved in the formation of pure subsets.
11. A method according to the preceding claim, wherein the basic training data are determined from high-resolution training data, - the high-resolution training data comprise high-resolution stress fields (THj i, TH,i, THjN) exerted on the turbine blade element (1' ; 1) and, for each of these fields, an associated label (k) identifying one of said typical responses (Mk), each high-resolution stress field (TH>i) giving values of said stress (T*H,i, Tju,i, Tj H,i) at different nodes of a high-resolution mesh of said blade element, - the method comprising a selection step (SO) of a restricted number (Q) of nodes of the high-resolution mesh, the selected nodes forming a relevant node group (S), - the basic stress fields (Tb T;, TN) basic training data being obtained by restricting the high-resolution constraint fields (TH, b TH i, TH,N) to said relevant node group.;
12. Method according to the preceding claim, wherein at least some of said relevant nodes are selected each on the basis of an approximate value of mutual redundancy information (Ij, j) between the node (j) under consideration and any of the other nodes (j') of the high-resolution mesh, said approximate value being determined on the basis of a spatial distance (djj) between the two nodes under consideration.
13. Electronic device comprising at least one processor and one memory, programmed to execute the numerical simulation method according to any one of claims 1 to 9 or the method for determining training data according to any one of claims 10 to 12.