Method for reducing the volume of training data for training a constraint field classifier, and associated electronic devices
By selecting relevant nodes based on mutual redundancy information and spatial distance, the method addresses the computational challenges of simulating turbine blade deformation and aging, reducing computation time and resources while maintaining accuracy.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- SAFRAN SA
- Filing Date
- 2020-12-16
- Publication Date
- 2026-05-15
AI Technical Summary
Numerical simulations of mechanical parts, particularly turbine blades in turbomachines, require extensive computing resources due to the complexity of simulating deformation and aging under varying stress fields, necessitating a large amount of computation time and resources, especially when characterizing the aging of turbine blades under uncertain temperature fields.
A method for determining training data for a constraint field classifier that selects a limited number of relevant nodes based on mutual redundancy information and spatial distance, reducing the size of training data while maintaining accuracy, using a simplified mechanical behavior model to simulate mechanical responses.
Significantly reduces computation time and resources required for simulating mechanical behavior of turbine blades by using a constraint field classifier trained on reduced data, maintaining simulation accuracy and efficiency.
Smart Images

Figure 00000040_0000 
Figure 00000040_0001 
Figure 00000040_0002
Abstract
Description
Title of the invention: Method for reducing the volume of training data intended for training a stress field classifier, and associated electronic devices. TECHNICAL FIELD OF THE INVENTION
[0001] In general, the invention relates to the field of numerical simulation of the mechanical behavior of mechanical parts, for example, their deformation and aging. It relates in particular to the application of artificial intelligence techniques during a preliminary simplification phase of the numerical simulation to be performed. The invention specifically relates to the determination of training data for training a stress field classifier, for example, a temperature field, this classifier enabling the identification, from a pre-established list, of a simplified mechanical behavior model well-suited to the stress field under consideration. TECHNOLOGICAL BACKGROUND OF THE INVENTION
[0002] Numerical simulation using finite element methods, which simulates the deformations undergone by a mechanical part subjected to various stresses, has made significant progress in recent years. Such simulations make it possible, in particular, to predict how a given mechanical part will age over a long period of use (for example, several years), which could not have been determined through direct testing or measurement. More generally, these simulations make the design and prototyping phase of a mechanical system faster and more efficient.
[0003] But some situations remain difficult to simulate, because they require a very large amount of computation time.
[0004] This is particularly relevant for characterizing the aging of a turbine blade in a turbomachine, for example, a high-pressure turbine located just downstream of the combustion chamber. In this case, a high, and generally inhomogeneous, temperature is exerted on the turbine blades, in addition to the pressure field and centrifugal force experienced by these blades.
[0005] For a given temperature field, exerted on this blade, simulating the deformation and aging of this or that element of the blade, for several cycles corresponding to several flights of an aircraft, requires very significant computing resources.
[0006] Furthermore, the temperature field exerted on such a turbine blade is not actually perfectly known. In order to be able to determine precisely a duration during which a blade will remain usable, and an uncertainty associated with this duration, it is necessary to therefore repeat the numerical simulation (already very demanding in itself) several times, for a whole (statistical) distribution of temperature fields.
[0007] With a classical finite element numerical simulation, the computation time or computing resources for such a characterization (number of processors or computing clusters to be made to work in parallel) then becomes prohibitive.
[0008] To reduce the computation time required to perform such simulations, an approach based on grouping different possible inputs (stresses) into clusters and a classification to recommend a reduced model (e.g., a reduced model of mechanical behavior) was proposed in the article "Localized Discrete Empirical Interpolation Method" by B. Peherstorfer et al., SIAM J. Sci. Comput. (2014). More recently, an approach based on such an order reduction of the mechanical problem and on a temperature field classifier was proposed in the following article: "Model order reduction assisted by deep neural networks (ROM-net)" by T. Daniel et al., Adv. Modell. and Simul. in Eng. Sci. (2020) 7:16.
[0009] According to this approach, instead of determining the mechanical deformation at each node of a mesh of the blade element to be characterized, with a number of degrees of freedom (scalars) which would therefore be equal to at least three times the number of nodes of the mesh, the mechanical response of the blade element is studied on the basis of a simplified mechanical behavior model, Mk, parameterized by a smaller number of degrees of freedom.
[0010] Such a mechanical behavior model Mk describes, for example, the deformation of the turbine blade element under consideration as a weighted sum of deformation modes typical of that element (for example, a pure bending mode, a mode corresponding to uniform stretching, another mode corresponding to stretching localized at one or the other end of the turbine blade element in question, etc.). The degrees of freedom, or in other words the parameters of this mechanical behavior model, then correspond, in this weighted sum, to the coefficients that weight the contributions of these different deformation modes.
[0011] The mechanical response of the turbine blade element, subjected to a given temperature field T, is then determined over time for several flight cycles by numerical simulation, based on this simplified model Mk. Thanks to the simplified nature of this model, it is possible to perform a simulation corresponding to a significant period of use, while maintaining a limited computation time.
[0012] Such a mechanical behavior model Mk is sometimes called a "reduced model" in this technical field. It is representative of the mechanical response of the element of a turbine blade (in particular, the deformation of this element) when this element is subjected to a given type of stress (for example, a given temperature field T). It is therefore, in a way, a typical response of the turbine blade element. Moreover, in the following, such a mechanical behavior model Mk is referred to interchangeably as a "reduced model" or a "typical response".
[0013] According to the approach proposed in the article by T. Daniel et al. mentioned above, before performing the numerical simulation in question, based on the reduced model Mk, this reduced model Mk is chosen according to the temperature field T exerted on the turbine blade element 1'. The reduced model Mk is selected, by a temperature field classifier 2', from among different reduced models Mb ... Mk, ..., MP, that is to say, from among different models of the mechanical behavior of the turbine blade element ([Fig. 1]). The reduced model Mk that is selected by the classifier 2' is, among the reduced models Mb ... Mk, ..., MP, the one that is most representative of a complete expected mechanical response for the turbine blade element 1' when it is subjected to the temperature field T in question.
[0014] The classifier 2', which includes for example a neural network, is trained using artificial intelligence techniques during a training step 20 ([Fig.2]). This training is carried out on the training data (TH>i, Mk), determined beforehand during a training data determination step 10.
[0015] This step 10 comprises a first step 11 of numerical simulation on a stress cycle of the turbine blade element ([Fig. 3]), corresponding, for example, to the stress experienced by this element during the entire flight of an aircraft powered by the turbomachine to which this blade element belongs. During step 11, the mechanical response of the turbine blade element is determined over time for this stress cycle, the blade element being subjected to a given temperature field TH i.
[0016] Such a numerical simulation is not necessarily sufficient to determine the long-term evolution of the mechanical characteristics of the blade element (which would require simulation over many successive cycles), nor its lifespan. However, it does allow for the identification of the type of mechanical behavior of this element when subjected to the temperature field TH i in question.
[0017] This comprehensive, but short-time, numerical simulation is performed for different so-called "high-resolution" temperature fields, TH.i, i=l,...,N. Each of these high-resolution temperature fields yields temperature values '-rJ 1 Hj to the different nodes j, j=l,...,QH of a mesh (so-called "high resolution" mesh) of the blade element.
[0018] This numerical simulation makes it possible to determine "complete" mechanical responses of the blade element, for these different temperature fields TH,i.
[0019] Then, in step 12, for each pair (TH>ib TH>i 2) grouping two of these high-resolution temperature fields, a dissimilarity of mechanical behavior ôn a is determined, or, in other words, an effective "distance" between the mechanical responses obtained for these two temperature fields TH>ib TH>i 2. This distance ôiU2 is, for example, representative of a difference between the two strain fields obtained respectively for these temperature fields. Different metrics are conceivable for determining this difference: it can be determined as being equal to the square root of the sum of the squares of the deviations at each node of the mesh, or as the sum of the absolute values of the deviations at each node of the mesh, or even as a distance in the sense of Grassmann, as is done in the aforementioned article by T. Daniel et al.
[0020] Next, in step 13, the various high-resolution mechanical responses determined in step 11 are grouped into several mechanical response groups Cb Ck, ..., CP, each grouping mechanical responses that are similar to each other, i.e., that exhibit, pairwise, a low dissimilarity δiU2 (a small "distance" between two elements of the same group). This step is generally called "clustering" in English (i.e., "data partitioning"), and the mechanical response groups Cb C2, CP are generally called "clusters".
[0021] Then, in step 14, for each cluster Ck, k=l.. .P, a reduced model Mk representative of the complete mechanical responses of the cluster considered (responses which are relatively similar to each other, by construction), is determined by an order reduction technique.
[0022] The training data ultimately obtained comprises various "labeled" data, each labeled data (TH.i, Mk) comprising: • one of the high-resolution temperature fields TH>i, i=l,.. ,,N, • as well as a label that identifies, among responses- types of blade element Mk , k=l.. .P, the standard response Mk corresponding to the temperature field TH>i considered (this standard response being the standard response associated with the cluster Ck in which is found the complete, high-resolution mechanical response obtained in response to the temperature field TH,i in question).
[0023] The label in question corresponds for example to the index k which identifies the standard response Mk, in the set of standard responses of the dawn element (this label is for example constituted by the index in question).
[0024] This set of labeled data (TH,i, Mk), i=l,.. .,N, is then used for training, i.e. for learning the classifier 2'.
[0025] It is important to note that the temperature fields TH i are grouped into different clusters Ck, each associated with a given typical response, not on the basis of a similarity between temperature fields, but on the basis of a similarity between the mechanical responses corresponding to these different temperature fields.
[0026] This distinction is particularly important because two temperature fields that appear close and similar to each other can in fact lead to very different mechanical responses, the interactions between thermal and mechanical loading being complex and dependent on the geometry of the turbine blade element.
[0027] Thus, two identical temperature fields, except for a very small portion of the blade element, can in fact lead to very different mechanical responses if the area in question is mechanically critical (for example, because it is thin, or subjected to high mechanical stress, or subject to stress concentrations). Conversely, two temperature fields that differ from each other over a large but less critical area, but which, on the other hand, take on the same values at the critical area in question, will probably lead to similar mechanical responses, even though these two temperature fields are apparently dissimilar.
[0028] The training data (TH>i, Mk), built on the basis of a similarity between responses, therefore allows a training of the classifier which is much more precise and reliable than with training data which would be based on a similarity between solicitations.
[0029] However, obtaining the training data (TH>i, Mk) requires significant computation time, since a numerical simulation (admittedly short-time, but on a complete mesh) is performed for each high-resolution temperature field TH,i. In practice, the number N of labeled data finally obtained is therefore limited.
[0030] For a mesh comprising approximately 40,000 nodes, obtaining just 500 labeled data points already requires a very significant computation time. And if one wants to obtain approximately 104 labeled data points, even with a reduced mesh comprising With only a few thousand nodes, it takes more than a day of computing by parallelizing the calculation on about fifty processors.
[0031] In this context, it would therefore be desirable to be able to obtain a larger number of reliable training data without significantly increasing computation time.
[0032] Various techniques for augmenting training data exist, particularly in the field of automatic image analysis. One known technique is based on applying various transformations and deformations to a given image. For example, an image representing a car is stretched, flipped, blurred, cropped, etc., to obtain additional images, which still represent a car.
[0033] But this type of technique is not applicable to the case considered here. Indeed, deforming one of the temperature fields THi would probably lead to a new temperature field giving a mechanical response completely different from that obtained with the initial temperature field TH,i (in other words, the label would not be preserved during such a transformation).
[0034] On the other hand, a difficulty arises from the size of each labeled data point itself. Indeed, to properly construct the Ck clusters and the associated typical responses Mk, given that the mechanical behavior of the blade element is initially unknown, the simulations mentioned above (short-time simulations) must be performed with a fine mesh. This is because no information is available in advance regarding the more or less critical nature of any given area from a thermomechanical point of view. The temperature field THi is therefore a high-resolution temperature field, and the resulting labeled data (TH>i, Mk) is thus large and cumbersome to process. Summary of the invention
[0035] In this context, a method is proposed for determining basic training data for training a constraint field classifier, • wherein high-resolution training data gather high-resolution stress fields exerted on a turbine blade element and, for each of these fields, a label associated with the stress field considered, each high-resolution stress field giving values of said stress at different nodes of a high-resolution mesh of said blade element, • the method comprising a step of selecting a limited number of nodes from the high-resolution mesh, the selected nodes forming a group of relevant nodes, • the basic training data, which includes basic constraint fields obtained by restricting high-resolution constraint fields to the relevant node group, • and in which at least some of said relevant nodes are selected each on the basis of an approximate value of mutual redundancy information between the node in question and any of the other nodes of said mesh, said approximate value being determined on the basis of a spatial distance between the two nodes in question.
[0036] The stress fields in question may, for example, be temperature fields or pressure fields exerted on the turbine blade element. For the sake of simplicity, the technical effects obtained using this method are presented below in the case of temperature fields. However, as will be apparent to those skilled in the art, this method can also be applied to other types of stresses exerted on the turbine blade element, such as pressure exerted on this element, for example.
[0037] Taking into account such mutual redundancy information when selecting relevant nodes makes it possible to remove a significant portion of the nodes while losing only a minimal portion of the information initially contained in the high-resolution temperature fields.
[0038] However, evaluating such mutual redundancy information is very computationally intensive due to the number of values to be determined, which is equal to the number of node pairs for the high-resolution mesh, a number that varies as the square of the number of nodes in the mesh. Calculating each of these mutual redundancy information values by direct statistical calculation would require a very significant amount of computation time.
[0039] Determining the different values of mutual redundancy information simply from the distance between the two nodes considered, in the form of a function of the distance between points, is significantly faster. Furthermore, this leads to an estimate, albeit approximate, but nevertheless relatively accurate, of the mutual redundancy information. Indeed, in practice, the mutual redundancy information proves to be well correlated with the distance between points (as illustrated by [Fig. 6], for example).
[0040] The fact that the mutual redundancy information is well correlated with the distance between points can be explained as follows. The temperature fields, or more generally the stress fields, used for numerical simulations of the mechanical behavior of turbine blades are generally at least partially random fields with Gaussian statistics. The temperatures at two nodes j and j', as given by this set of fields, can then be seen as Gaussian random vectors. And it can be mathematically shown that the mutual redundancy information / (x1, X2) between two such vectors X1 and X2 is expressed as:
[0041] j(x\ X2) = 4 ln(l-p2)
[0042] where p is the correlation between X1 and X2.
[0043] Moreover, these random temperature fields generally have isotropic correlation functions. The correlation p then depends only on the distance between the nodes considered, which ultimately explains why the mutual redundancy information can be expressed, at least approximately, as a function of distance alone.
[0044] The link between mutual redundancy information and the geometric distance between points is therefore a consequence of the type of temperature fields used (Gaussian, and with an isotropic correlation function). It should be noted, however, that this causal relationship is actually quite indirect. Indeed, the temperature fields are generally obtained by imposing surface temperature fields of known statistics on the surface of the blade element, and then taking into account the propagation of heat within the turbine blade element to obtain the temperature at every point. The fact that the volumetric temperature fields thus obtained ultimately have statistics leading to a good correlation between mutual redundancy information and the geometric distance between points is therefore far from straightforward at first glance.
[0045] The method just presented is implemented by a computer or any computer system capable of executing logical operations.
[0046] In addition to the characteristics mentioned above, the method just presented may have one or more of the following complementary characteristics, considered individually or in all technically possible combinations: • said stress exerted on the turbine blade element is a temperature exerted on that element, the high-resolution stress fields and the basic stress fields being respectively high-resolution temperature fields and basic temperature fields; • The mutual redundancy information between the two nodes considered is representative of a level of information redundancy between: • the values of said constraint at the first of the two nodes, as given by the set of high-resolution constraint fields, and • the values of said constraint at the second of the two nodes, as given by the set of high-resolution constraint fields; The method includes a calibration step which comprises the following steps: • for each pair of nodes in a subset grouping a portion of the nodes in the high-resolution mesh: • determination of a reference value for the mutual redundancy information between the two nodes considered, by statistical calculation based on the values of said constraint, at the level of these two nodes, as given by the various high-resolution constraint fields, and • Calculation of the spatial distance separating the two nodes considered, • adjusting a curve so that it is representative of a link between said reference values of mutual redundancy information, and the corresponding spatial distances; the approximate value of the mutual redundancy information between two nodes of the network is the value given by said curve, for the spatial distance separating the two nodes considered; for each high-resolution stress field, the label associated with the high-resolution stress field in question identifies a typical response, representative of an expected mechanical response for said blade element in response to said high-resolution stress field, among different typical responses of said turbine blade element; each relevant node is further selected on the basis of a degree of relevance of the value of said stress at the node considered, with respect to the mechanical behavior of said blade element, said degree of relevance being representative of a level of dependence between, on the one hand, the label associated with the high-resolution stress field considered, and, on the other hand, the value of said stress at the node considered for this stress field; The relevant node group initially contains a first node, which is the node in the high-resolution mesh with the highest degree of relevance. An addition step is then performed, during which an additional node is added to the relevant node group. This additional node is selected from among the nodes in the mesh that do not... not yet part of said group, so as to maximize an equal quantity: to the degree of relevance of the node considered, less an average level of redundancy between this node and the nodes already present in the group of relevant nodes, the addition step then being executed again several times; The average level of redundancy between the node in question and the nodes already present in the relevant node group is equal to the arithmetic mean of the approximate values of the mutual redundancy information between said node and the nodes already present in the relevant node group. The additional node added to the relevant node group is the node that maximizes the following quantity: r(T1) = 1 / V(r1) where i Card^S) s1 ( ) is the index identifying the node in question, in the high-resolution mesh, and where C&rd(S) is the number of nodes contained in the relevant node group during the current iteration of the addition step; The method further includes a data augmentation step in which additional training data is determined from said basic training data, each basic constraint field being labeled with the corresponding high-resolution constraint field, the data augmentation step comprising the following steps: • form at least one pure subset, which groups together basic stress fields associated with the same standard response of the turbine blade element, and which has a convex envelope containing no basic stress field associated with another standard response of said element, • for at least one of said pure subsets, determination of one or more additional training data, each additional training data comprising an additional constraint field, obtained by linear combination of basic constraint fields of the pure subset considered, assigned positive weighting coefficients and normalized to 1, the additional constraint field being accompanied by a label which is the same as for the basic constraint fields of said combination; The step of forming one or more pure subsets includes, for each of these subsets, a purity test step of the subset in question, during which it is verified that no basic constraint field associated with a standard response different from the common standard response is found. to the basic stress fields of the considered subset, is not located inside the convex envelope of said subset; • Each basic stress field groups a set of values of said stress at different nodes of a given mesh comprising a total number of nodes Q, and for each basic stress field, an associated stress subfield is obtained by projecting the basic stress field onto a subspace with a number of degrees of freedom less than Q, and during the purity test step: • a projection test is performed, during which it is determined whether, for the stress subfields associated with the basic stress fields of the subset to be tested, a convex hull of the set grouping these stress subfields contains no stress subfield associated with one of the basic stress fields for which the standard response is different from the standard response common to the basic stress fields of the subset to be tested, • and, when the result of this projection test is positive, it is determined that the subset of stress fields considered is pure.
[0047] The invention also relates to a numerical simulation method of the mechanical response of a turbine blade element subjected to a given stress field TR comprising: • a preliminary training phase including: • a determination of basic training data, or a determination of basic training data and additional training data in accordance with the claim, explained above, • training a constraint field classifier based on said basic training data, or based on said basic training data and said additional training data, and • a usage phase comprising the following steps: • identification of one of the standard responses for said blade element by the constraint field classifier receiving said constraint field as input, • and numerical simulation of the mechanical response of said blade element, subjected to said stress field, in accordance with the model simplified mechanical behavior defined by the standard response previously identified by the classifier.
[0048] The invention also relates to an electronic device comprising at least one processor and one memory, programmed to determine basic learning data, or to determine basic learning data and additional learning data in accordance with the method described above.
[0049] The invention also relates to an electronic device comprising at least one processor and one memory, programmed to execute the numerical simulation method which has just been presented.
[0050] The invention and its various applications will be better understood by reading the following description and examining the accompanying figures. BRIEF DESCRIPTION OF THE FIGURES
[0051] The figures are presented for illustrative purposes only and are in no way limiting of the invention.
[0052] [Fig.1] Fig.1 schematically represents an operation of identification, by a temperature field classifier, of a typical response adapted to a given temperature field, exerted on a turbine blade element.
[0053] [Fig.2] Fig.2 schematically represents a step in determining training data and a training step of a classifier such as that of [Fig.1].
[0054] [Fig. 3] [Fig. 3] schematically represents steps in the determination step of training data from [Fig.2].
[0055] [Fig.4] Fig.4 schematically represents a method for determining training and learning data of a classifier such as that of [Fig.1].
[0056] [Fig. 5] [Fig. 5] Schematically represents, in perspective, an example part mechanical for which the process of [Fig.4] is implemented.
[0057] [Fig.6] Fig.6 schematically represents mutual information values of redundancy between nodes of the mesh of [Fig.5].
[0058] [Fig.7] Fig.7 schematically represents steps for selecting relevant nodes, among the nodes of a complete mesh such as that of [Fig.5],
[0059] [Fig.8] [Fig.8] schematically represents the part of [Fig.5], in perspective, with the set of relevant nodes selected through the steps of [Fig.7].
[0060] [Fig.9] Fig.9 schematically represents the steps of a process data increase.
[0061] [Fig. 10] The [Fig. 10] is a schematic representation, in a reduced two-dimensional space, of two clusters of temperature fields, and of a pure subset identified in one of these clusters.
[0062] [Fig. 11] The [Fig. 11] is a schematic representation, in a reduced two-dimensional space, of different temperature fields each reduced to a point, these temperature fields belonging to different clusters.
[0063] [Fig. 12] The [Fig. 12] schematically represents steps implemented to form a pure subset, such as that of the [Fig. 10].
[0064] [Fig. 13] The [Fig. 13] schematically represents steps implemented to test the purity of such a subset. DETAILED DESCRIPTION
[0065] The embodiment described below corresponds to a case in which the stress exerted on the turbine blade element considered is a temperature exerted on that element (at different points on that element). In other embodiments, the stress in question could be of another type. It could, for example, be a pressure exerted on that blade element by a fluid flowing in the turbine.
[0066] Figure 4 represents a method for determining training data and training a temperature field classifier, comparable to that described in the section relating to the technological background (with reference to Figures 1 to 3) but also incorporating: • a variable selection step S0, and • a step A0 of data increase.
[0067] Step 10 of [Fig.4] is the same as that described previously, with reference to Figures 2 and 3. It allows us to determine a first set of training data (TH>i,Mk) bringing together the high-resolution temperature fields TH>i mentioned above, and, associated with each of these temperature fields, the label that identifies the corresponding typical response Mk (i.e. the typical response Mk corresponding to the cluster Ck to which this temperature field belongs).
[0068] The variable selection step S0 is a step in which, for each high-resolution temperature field TH.i, a corresponding basic temperature field T is determined; the basic temperature field T' gives temperature values at different nodes of a simplified mesh that comprises fewer nodes than the high-resolution mesh. This simplified mesh consists of a group of relevant nodes S, selected from among the different nodes of the high-resolution mesh. This simplified mesh may, for example, comprise a number Q of nodes 100 times, or even 400 times smaller than the number of QH nodes in the high-resolution mesh.
[0069] The simplified mesh is determined here so that the basic temperature fields T;, significantly simplified compared to the high-resolution temperature fields TH i, are nevertheless as relevant as possible from the point of view of the thermomechanical response of the turbine blade element 1' ; 1 whose behavior we want to simulate.
[0070] Step S0 thus makes it possible to obtain a basic training dataset (T;, Mk), which brings together the basic temperature fields T; in question, each with its label, this label being the same as for the high-resolution temperature field TH>1 which has been "compressed" in the form of the basic temperature field T; considered.
[0071] Next, during the data augmentation step A0, additional training data (T'i, Mk) is determined from the basic training data (T, Mk). The additional training data (T'i, Mk) consists of additional temperature fields T'i, each with a label that identifies one of the typical responses Mb...Mk,...,MP of the turbine blade element.
[0072] Next, in step 20, the temperature field classifier is trained with the basic training data (T;, Mk) and the additional training data (T'i, Mk) obtained previously.
[0073] In the following, the variable selection step S0 is described in more detail, with reference to Figures 5 to 8. The data augmentation step A0 is then presented, with reference to Figures 9 to 13.
[0074] Step S0: variable selection
[0075] Figure 5 shows, in perspective, the mechanical part 1, which serves here as an example to illustrate the implementation of this method for determining training data. This figure shows the high-resolution mesh, with QH nodes, mentioned above, as well as the positions of 2 and three nodes of this mesh. In this example, the number QH of nodes in the high-resolution mesh is equal to 42445.
[0076] Inset (c) of [Fig. 8] shows, in the form of thick dots, the positions of the different nodes of the relevant node group S obtained for this same example, using the variable selection technique in question. In this example, the number Q of nodes in the relevant node group S is equal to 87.
[0077] Here, the relevant nodes are selected: • based on mutual redundancy information ["Centre nodes, and based on a degree of relevance y) of each node, with respect to the assignment of such and such a typical response Mk to a given temperature field.
[0078]
[0079]
[0080] The mutual redundancy information Ijr is representative of a level of information redundancy between the temperature value at a first node j of the high-resolution mesh, and the temperature value at a second node j' of this mesh, for the temperature field distribution corresponding to the set of high-resolution temperature fields. The relevant nodes are selected so as to have a high degree of relevance, while being little redundant with the nodes already selected, which are already part of the relevant node group S. To do this, we form the relevant node group S as follows. First, we integrate into the relevant node group S a first node, which is the node j of the high-resolution mesh having the highest degree of relevance j ( Then we It then executes, several times successively, a node addition step. During this step, an additional node, which is selected, is added to the relevant group of nodes S: among the nodes of the high-resolution mesh that are not yet part of said group, and so as to maximize a quantity equal to the degree of relevance T ( •LPX / of the node i considered, less an average level of redundancy between this node and the nodes already present in the relevant node group S.
[0081] More precisely, the additional node, added to the group of relevant nodes S during the addition step, is the node i that maximizes the following quantity:
[0082] t (r1}-___1__ y T ' Card(s) S11 J
[0083] where i is the index identifying the node in question, in the high-resolution mesh, and where Card(S) is the number of nodes already contained in the relevant node group S during the current iteration of the addition step.
[0084] Here, the iterations of the node addition step cease when the above quantity no longer changes significantly over several successive iterations. For example, if this quantity has not changed by more than 5% (in relative value) over the last 5 iterations, this addition step is no longer repeated.
[0085] Alternatively, this addition step could also be discontinued when the quantity in question no longer decreases significantly from one iteration of the step to the next. It could also be planned to stop the iterations when the quantity in question falls below a given threshold, or when a predetermined number of iterations is reached.
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094] reached. We could also plan to stop the iterations when the relevance information j ( of the node to be added) falls below a given threshold. This criterion is interesting because the quantity to be maximized at each iteration (quantity above) can take a high value for a node with low relevance information j (and therefore not very relevant) but also with low redundancy with the already selected nodes. As already mentioned, the mutual redundancy information Ijr represents the level of information redundancy between the temperature value Tj at a first node j of the high-resolution mesh, and the temperature value Tj- at a second node j' of this mesh. More precisely, this mutual redundancy information, which reflects a greater or lesser degree of interdependence between the temperatures Tj and Tj, is defined by the following relationship: ..........' dtj dtj' p / fl PjW') / Or : - the integration domain is R2, - n ( f-j1 is the probability density, for the temperature field distribution given by the set of high-resolution temperature fields TH4, ..., THN, to have a temperature tj at node j, - p is the probability density, for this same distribution, of having a temperature tj at node j', and where - p(fi') is the (joint) probability density function, for this same J d distribution, to have a temperature tj at node j and a temperature tj at node j'. As for the degree of relevance j (jdj of node j), with respect to the assignment of such and such a typical response Mk to a given temperature field, it is representative of a level of dependence between: • on the one hand, the label number (i.e., the index k identifying the appropriate standard response), associated with a given high-resolution temperature field, and, • on the other hand, the temperature value at node j, given by this temperature field. The degree of relevance j ( is a particular form of mutual information redundancy, in the case where a first variable is the temperature Tj at node j, while the second variable, Y, has a discrete value and represents the number of
[0095] the label associated with temperature field ■ = l(TJ y) with i(tj, y) = dtJ
[0096] where:
[0097] - the integration domain is R,
[0098] - p(t^) is the probability density function for the temperature field distribution given by the set of high-resolution temperature fields Th,i, ..., THN, to have a temperature tJ at node j,
[0099] - is the probability for the standard response to be Mk
[0100] -p ( H|l<) is the probability density for the distribution of fields of temperature given by the set of high-resolution temperature fields Th,i, ..., Th,N, to have a temperature tj at node j knowing that the standard response is Mk, and where
[0101] - pj is the probability density function for the distribution of fields of temperature given by the set of high-resolution temperature fields Th,i, ..., TH,n, to have a temperature tJ at node j and for the standard response to be Mk, simultaneously.
[0102] The degree of relevance j ( • ct the mutual information of redundancy Ijr can to be determined, on the basis of the above expressions, after a prior estimation of the different probability functions involved in these expressions.
[0103] These probability functions can for example be estimated, for the (discrete) distribution of temperature fields and labels constituted by the set of high resolution temperature fields Th,i, ..., THjN, with their labels, by "binning" techniques (i.e. by sampling techniques).
[0104] By way of example, the following article describes such a "binning" technique, which can be used to calculate such mutual redundancy information from discrete data: "Estimating mutual information", A. Kraskov, H. Stôgbauer, and P. Grassberger, Phys. Rev. E 69, 066138 - 23 June 2004; Erratum: Phys. Rev. E 83, 019903 (2011).
[0105] In any case, a determination of the mutual redundancy information Ij j-between two nodes j and j', by such a statistical calculation based on the temperature values TjH,i, ..., TjH,N, and Tj h,i, ..., Tj h,n at the level of these two nodes, is computationally intensive.
[0106]
[0107]
[0108]
[0109]
[0110] And the number of terms Ijr to be calculated is very large since it represents the number of pairs grouping any two nodes j, j' of the high-resolution mesh. In the numerical example corresponding to [Fig. 5], the number of pairs is on the order of 10⁹, for example. Also, to reduce the computation time required to evaluate the mutual redundancy information Ijj between nodes, rather than determining it by direct statistical calculation (which would be very long), we determine here an approximate value J ( djj>^ of this quantity from the spatial distance, djr, between the two nodes j and j' considered. In other words, we consider here that the mutual redundancy information Ijj- can be expressed, to a good approximation, as a function of the distance djj- in question: Ijj^lÇdjj'y. This distance djj- is equal to the geometric distance II II separating the two nodes j, j' in question. This distance II JJ 2 is quick to calculate, from a numerical point of view. Evaluating the mutual redundancy information Ij r in this way allows the relevant node group S to be formed much faster than by evaluating all the Ijj- terms by direct statistical calculation. The one-variable function ï, which is used to calculate the approximate values J^d jj1} from the distances djj-, is determined beforehand, during a calibration step SI 1 (figure 7), by determining the mutual redundancy information Ijj- in an "exact" way, that is, by direct statistical calculation, for some nodes of the high-resolution mesh, in order to calibrate the curve, or, in other words, the function I. The SU calibration process includes the following steps: • S12: for each pair of nodes (j, j') of a subset grouping a portion of the nodes of the high-resolution mesh: • determination of a reference value j (pi pJ J (the "exact" value) of the mutual redundancy information Ijj- between the two nodes j and j' considered, by statistical calculation based on the temperature values TjH> b ..., TjH> N, Tj H> b ..., Tj H> N, at the level of these two nodes, as given by the different high-resolution temperature fields TH.i, ..., THjn, and • calculation of the spatial distance djj separating the two nodes j and j' considered, then [YES]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120] • S13: adjustment of curve I so that it is representative of a link between the reference values j ( y J ) a^ns^ obtained, and the corresponding spatial distances. The subset of nodes, for which mutual redundancy information is calculated in an "exact" way, may include a number of nodes ranging, for example, from one thousandth to one twentieth of the total number QH of nodes in the high-resolution mesh. The curve î is a curve parameterized by one or more parameters which are adjusted (for example by a least squares method) so that this curve is as representative as possible of the reference data Figure 6 shows reference values jy J y / ') of mutual information of redundancy Ijj- between nodes, obtained for part 1 in Figure 5. In this figure, each point represents a pair (djj', jy J y J') ). This figure also shows the curve I obtained by fitting these reference data. In this example, approximately 750 reference values j [ yJ yJ'J were determined to calibrate the curve î. In this case, the curve î represents the following single-variable function: ï(r)=Iæ+ y1.(r1-r)ai.H(r1-r) + y2.(r2-r)ff2.H(r2-r) where H is the Heaviside stair-step function, and where Ix, yy Cp Ct py -^2' Ct2 are the parameters adjusted during step S13. As can be seen in this example, the mutual redundancy information Ijj' is well correlated with the distance djj-, and can indeed be expressed, to a good approximation, in the form of a function I of this distance. This property, exploited here in a particularly interesting way, can be explained as follows. The temperature fields used for numerical simulations of the mechanical behavior of turbine blades are generally at least partially random fields with Gaussian statistics (this is notably the case for the high-resolution temperature fields used here). The temperatures at two nodes j and j', as given by this set of fields, can then be seen as Gaussian random vectors. And it can be mathematically shown that the mutual redundancy information / (x1, X2) between two such vectors X1 and X2 is expressed as: x2)= 4 ln(l-p2) where p is the correlation between X1 and X2.
[0121] Moreover, these random temperature fields generally have isotropic correlation functions. The correlation p then depends only on the distance djr between the nodes considered, which finally explains why the mutual redundancy information Ijj- can be expressed as a function of the single distance d^.
[0122] In any case, at the end of the calibration step SI 1, the curve I is completely determined, to then be used for the selection of the relevant nodes.
[0123] The relevance degree jj is calculated during a step S14, for all nodes of the high-resolution mesh. It is determined by direct statistical calculation. The computation time required to perform this step remains reasonable, however, because, unlike the case of mutual information, it involves calculating a value per node (and therefore QH different values), and not a value per pair of nodes (i.e., Qh.(Qh- l) / 2 different values).
[0124] Then, in a step S15, the relevant node group S is determined, as explained above, by selecting relevant nodes based on the degree of relevance and mutual redundancy information. As the relevant node group S grows progressively, the additional node added to the nodes already present in this group is node i that maximizes the following quantity:
[0125] T (t*}-___1__ V 1 f Card(S) )
[0126] As already indicated, inset (c) of [Fig. 8] shows the different nodes of the relevant node group S obtained for the high-resolution mesh example of [Fig. 5]. In this example, the node addition iterations were stopped once the above quantity ceased to vary significantly from one iteration to the next. 87 relevant nodes were thus selected (which shows, incidentally, that significant compression of the initial high-resolution mesh was possible). 303 seconds were required to identify this group of 87 relevant nodes S (i.e., to perform 86 successive executions of the addition step). By comparison, when the mutual redundancy information Ijr is evaluated by direct statistical calculation, it takes, on the same computer, 6469 seconds to execute only the first 6 iterations of this selection algorithm (and thus select 7 relevant nodes).This illustrates the very significant acceleration made possible by the approximate calculation technique for mutual information presented here.
[0127] Inset (a) of [Fig. 8] shows a group _S' of 87 nodes, selected from the high-resolution mesh using a different selection technique. In this case, the nodes in this group were selected based on a purely geometric criterion, so as to obtain, for the selected nodes, a substantially homogeneous node density over the entire part (for example, so that the volume of a zone specific to each selected node, centered on that node and not containing any other relevant node, varies little from one selected node to another). This selection is made without taking into account the mechanical responses of the part.
[0128] Inset (b) of [Fig. 8] shows a group _S” of 87 nodes, selected from the high-resolution mesh using yet another selection technique, based on a univariate filter. As can be seen in this figure, the nodes selected in this way are concentrated mainly at a central constriction in part 1, and on the surface of this part (an area that is a priori quite critical from a thermomechanical point of view). On the other hand, it can be seen that virtually no nodes are selected elsewhere than in this area, and that many nodes are located (probably unnecessarily) in the same place.
[0129] The selection technique presented above, based on a degree of relevance corrected for redundancy, leads to a distribution of nodes (inset (c)), which, compared to the homogeneous distribution of inset (a), is denser at the central narrowing, while covering the whole of this piece (unlike selection by univariate filter).
[0130] The overall relevance degree D of the group of relevant nodes S, S' and S”, and its overall redundancy level R are given below, for these three selection techniques: • (a): purely geometric distribution; D=0.009 and R=0.1072; • (b): selection by univariate filter; D=0.0671 and R=0.8129; • (c): maximizing relevance while minimizing redundancy; D=0.046 and R=0.111.
[0131] A purely geometric distribution therefore leads to a low redundancy R, but also to a low relevance D.
[0132] Selection by univariate filter leads to high relevance D, but to high redundancy R.
[0133] As for selection by maximizing relevance and minimizing redundancy, it leads to both high relevance D and low redundancy R.
[0134] The overall relevance degree D of a group of relevant nodes given above is defined by the following formula:
[0135] 0 =
[0136] The overall redundancy level R of a group of relevant nodes is defined by the following formula: 101371
[0138] For the method of determining training data described here, the selection technique, used during the step of selecting variables S0, is technique (c), maximizing relevance while minimizing redundancy.
[0139] Alternatively, another selection technique, for example by univariate filter or by purely geometric distribution, could however be used during step S0.
[0140] In any case, at the end of the variable selection step, we have a set of basic temperature fields Tb.. .,T;,.. ,,TN, with their labels, less voluminous than the initial high-resolution temperature fields Th,i, ... ,TH>i,... ,Th n.
[0141] In the next step A0, additional training data are determined, by data augmentation, from these basic training data.
[0142] And ape A0: data augmentation
[0143] The data augmentation step A0 comprises the following steps ([Fig.9]): • Step A1: form different subsets pure Si, and • Step A2: within each pure subset, determine data additional learning through convex linear combinations.
[0144] At step Al, each pure subset Si is formed in such a way: • to group together basic temperature fields Tj that are associated with the same typical response Mk of the turbine blade element (i.e., that belong to the same cluster Ck), • and so as to have a convex envelope Ei which does not contain any basic temperature field Tj which is associated with another response-type Mk-, kVk, of the turbine blade element (i.e. no basic temperature field T; foreign, belonging to a cluster other than the Ck cluster).
[0145] Then, during step A2, for each subset Si, different additional training data are determined, each additional training data comprising: • one of the additional temperature fields T'; in question, which is obtained by linear combination of the basic temperature fields Tj of the pure subset Si considered, assigned positive weighting coefficients and normalized to 1, and • a label (for example, the index k identifying the standard response Mk), which is the same as for the basic temperature fields Tj that were used to obtain this additional temperature field T';.
[0146] Step A1, of forming the pure subsets Sb 1=1..M, is now described in more detail. This step comprises the following steps: • Step A10: among the base temperature fields T, belonging to the same cluster Ck, selection of "seeds" SDb 1=1.. .M, for the subsequent growth of pure subsets Sb • Step Al 1: for each seed SDb, formation of a subset Sb initially containing said seed, and which is grown until its purity is lost.
[0147] The A10 seed selection step is presented first, and the Garlic growth step, which includes a specific purity test, is described next.
[0148] Step A10: Seed selection
[0149] Each SDi seed consists of one of the basic temperature fields Ti belonging to the considered cluster Ck, judiciously selected.
[0150] Here, the SDi seeds are chosen, within the considered cluster Ck, so as to be somewhat distant from the cluster boundaries. More precisely, they are chosen so as to be a minimum distance from the basic temperature fields belonging to the other clusters Ct, k'k. This distance is a distance in the sense of the distance described above (in the part relating to the general technological background), that is to say, a distance which reflects a dissimilarity of the mechanical responses, obtained respectively for the two basic temperature fields Tj, Tj considered (dissimilarity of mechanical behaviors).
[0151] Thus, in each cluster Ck, the SDi seeds are selected so as to exhibit, with respect to the different basic temperature fields T; • of the other clusters Ckb kVk, a dissimilarity ôy- of mechanical response greater than a given threshold ôi.
[0152] In each cluster, it is advantageous to choose seeds far from the base temperature fields of other clusters Ckb because they are conducive to the subsequent growth of a pure subset Sb. Indeed, thanks to this distance, this subset will be able to grow a minimum in size before a base temperature field, belonging to another cluster Ck-, is found in its convex hull Ei.
[0153] More generally, this separation limits the risk that the convex hull Ei of the subset Si will encompass, in the temperature field space, a part of another Cluster Ckb, whether or not that part contains a basic temperature field. In other words, for the example in [Fig. 10], choosing seeds in cluster Ci that are far from the crosses representing the basic temperature fields of cluster C2 prevents a subset from one of these seeds, does not encompass a part of the domain corresponding to cluster C2, a domain which is represented in white on the figure, whether this part contains a cross or not.
[0154] Furthermore, within each cluster Ck, the seeds are chosen so as not to be too isolated from the other basic temperature fields of that same cluster, so that the seed is conducive to the subsequent growth of a subset. More precisely, each seed is chosen so that at least one other basic temperature field of the same cluster Ck is located at a "distance" b,, from the seed that is less than a given threshold b2 (in other words, which exhibits, with the seed in question, a dissimilarity b,, of mechanical response lower than the threshold b2 in question).
[0155] To obtain seeds that are not too isolated from the other base temperature fields in their cluster Ck, and that do not have close neighbors belonging to a foreign cluster Ck, we begin by selecting, in the form of a list Fk, all the base temperature fields that satisfy these two criteria in the cluster Ck under consideration. This list is obtained by removing (deleting) from the complete list grouping all the base temperature fields of the cluster Ck under consideration: • all basic temperature fields 1), having a neighbor Ti' which belongs to another cluster and which, with respect to T;, is located at a distance by- less than the threshold ôb and • all base temperature fields T; which have no neighbor Tr-belonging to the same cluster and located at a distance by • from T; less than the threshold b2
[0156] The thresholds ôi and b2 are defined here as being equal to Ei.bref and e2.ôref respectively, where ôref is a distance representing a dimension of the cluster Ck under consideration (the values of the thresholds used may therefore vary from one cluster to another). The coefficients Ei and e2 are less than 1 (and positive). For example, Ei can be between 0.01 and 0.05, and e2 can be between 0.1 and 0.3 (this means that we also select seeds that are relatively "isolated." This is useful because the data augmentation also aims to generate data in areas that lack it, i.e., those that are poorly represented by the initial data).
[0157] ôref is, for example, determined to be the distance between: • a base temperature field of the cluster Ck, considered to be representative of that cluster, and • the base temperature field, foreign to the Ck cluster, and which is closest to the representative of Ck.
[0158] The representative of the Ck cluster is, for example, the basic temperature field whose mechanical response is closest to an average of the mechanical responses obtained for the different basic temperature fields of this cluster.
[0159] Once the list Fk of temperature fields in the cluster Ck that are neither too isolated nor too close to the boundaries has been obtained, a given number of seeds are selected from this list. This number is, for example, previously set by a user.
[0160] The seeds can be selected, from the Fk list, so as to be distributed evenly in the cluster.
[0161] Here, the seeds can be selected from the Fk list as follows. The first seed, SDi, selected is the base temperature field that represents the cluster Ck under consideration. The other seeds are selected one after the other, such that each new seed selected is the base temperature field in the Fk list that is furthest from the nearest seed. The second seed, SD2, for example, is the base temperature field in the Fk list that is furthest from the first seed, SDi. And the third seed, SD3, is the base temperature field in the Fk list that is furthest from either the first seed, SDi, or the second seed, SD2, depending on whether the nearest seed to SD3 is SDi or SD2.
[0162] Thus, if Lk denotes the current list of selected seeds, the next seed added to this list is the base temperature field Tj whose index j is given by the following formula:
[0163] j = arg( max [minô ^\a£Fk\Lk[b&Lk
[0164] where ô is the distance between temperature fields mentioned above (also called mechanical response dissimilarity).
[0165] Selecting the seeds in this way allows for homogeneous coverage of the entire Ck cluster, and therefore, subsequently, the creation of pure subsets that are also homogeneously distributed within the cluster. This is useful because pure subsets distributed inhomogeneously would lead, after the creation of additional training data within these subsets, to an overrepresentation of certain areas of the cluster, which is not optimal for the classifier's training.
[0166] Step A1: growth of pure subsets.
[0167] Once the SDi seeds, 1=1..M, have been selected, for each SDi seed, a corresponding subset Si is grown from the seed in question. ab
[0168] Initially, the subset Si contains only this seed SDb. It is then grown by performing the following steps ([Fig. 12]): • A10: adding one of the basic temperature fields T to the subset Si, the added temperature field having the same response type Mk as the other temperature fields in the subset Si (in other words, the added temperature field belongs to the same cluster Ck as the other elements of SJ, then • Al 11: purity test, for the subset Si thus augmented.
[0169] If step Al 11 shows that the subset Si, thus augmented by another basic temperature field, is still pure, step Al 10 is executed again, then step Al 11, and so on until step Al 11 shows that the subset Si is no longer pure. The iterations of steps Al 10 and Al 11 are then stopped, and the last basic temperature field to have been added to the subset is removed from this subset Si, thus making it pure again.
[0170] The growth of the subset Si is therefore carried out step by step. More precisely, during the addition step Al 10, the base temperature field T; which is closest to the seed SDi (in the sense of distances ôÿ), is added to the subset Si, among the base temperature fields of the cluster Ck which have not yet been selected (i.e., which are not already part of the subset SJ.
[0171] It should be noted that the subsets Si are thus grown step by step, based on a particular metric ôÿ which represents the dissimilarity between the mechanical responses associated with the temperature fields, and not based on a usual (e.g., Euclidean) distance between the temperature fields themselves. This allows each subset to grow with a good probability of remaining pure (since the purity criterion in question concerns the mechanical responses associated with the temperature fields). In other words, growing the subsets Si based on the same metric ôÿ as that initially used to identify the different clusters Ci, C2, ... CP (i.e., the different classes) of temperature fields optimizes the probability that the subset remains pure during its growth. This therefore makes it possible to obtain larger pure subsets.This precaution is welcome in practice, because even when growing the pure subsets in this way, their maximum size generally remains small (they typically contain only a few temperature fields, on the order of 5 to 10, on average, for the examples presented here).
[0172] Similarly, the fact of pre-selecting the seeds (in step A10) on the basis of this same metric ôÿ, initially used to identify the different clusters Ci, C2, ... Cp, makes these seeds good candidates for the subsequent growth of pure subsets.
[0173] Purity test
[0174] During the purity test step Al 11, it is verified, for the subset Si to be tested, that no basic temperature field Tr, associated with a standard response Mk- different from the standard response Mk common to the basic temperature fields T of the subset Si considered, is located inside the convex hull Ei of this subset Sp
[0175] In other words, it is verified that no basic temperature field, belonging to a cluster other than the cluster Ck of the subset to be tested, is located inside the convex envelope Ei of the subset Sp. When this condition is verified, the subset in question is considered pure.
[0176] This test amounts to verifying, for each basic temperature field Tr belonging to another cluster Ck-, that this temperature field Tr cannot be expressed as a linear combination of the basic temperature fields Ti belonging to the subset Si in question, this linear combination having positive coefficients and normalized to 1.
[0177] The criterion to be tested is therefore relatively computationally intensive, even if the number of degrees of freedom of the basic temperature fields T; is relatively reduced, compared to the initial high-resolution temperature fields TH i (thanks to the prior step of variable selection S0).
[0178] Also, to reduce the computation time required to perform the purity test Al 11, the purity of each subset Si is tested by projection, i.e. on the basis of temperature subfields Tp obtained by projection of the basic temperature fields Tj onto a subspace having a number of degrees of freedom less than the number Q of nodes of the mesh on which the basic temperature fields T are defined.
[0179] As indicated in the description of the variable selection step S0, each base temperature field T; groups a set of Q temperature values T / ,.. ..,T;q , which are the temperature values of that field at the different nodes of the mesh in question (which, in this case, is an already simplified mesh, compared to the initial high-resolution mesh which has QH nodes).
[0180] The subspace onto which the basic temperature fields T are projected is a space of dimension QR, that is, having a number of scalar degrees of freedom QR less than Q (in practice, it may be a subspace of very small dimensions, for example, only 5 dimensions instead of 87 or 42445 dimensions). This subspace corresponds here to the set of fields of temperature that gives temperature values in a number QR of mesh nodes only, instead of giving these values in all mesh nodes on which the basic temperature fields T; are defined.
[0181] Here, each temperature subfield Ti thus gives temperature values at different nodes of a restricted node group, comprising only QR nodes (for example, only 5 nodes) instead of Q nodes. These nodes are chosen from among the nodes of the (reduced) mesh on which the basic temperature fields T are defined; that is, from among the Q nodes of the relevant node group S determined in step S0. In this case, they are chosen randomly from among the Q nodes of the relevant node group S.
[0182] Each temperature subfield T i is then obtained, from the corresponding basic temperature field T;, keeping, among the temperature values T;1,.. ..,TiQ given by Ti, a restricted number of these temperature values, which are those at the level of the nodes of the restricted group in question.
[0183] It should be noted, however, that other ways of projecting the temperature fields T; are conceivable. Any projection onto a subspace of RQ is conceivable. This subspace may, for example, correspond to the QR first components of the canonical basis of RQ, or to the QR first left singular vectors in the singular value decomposition (SVD) of the collection of temperature vectors T / ,.. ..,TiQ at the nodes of the reduced mesh..
[0184] In any case, carrying out the Al 11 purity test mentioned above, not on the basis of the basic temperature fields T;, but on the basis of the less voluminous subfields T p, makes it possible to accelerate this purity test step.
[0185] Moreover, surprisingly, when the result of this purity test by projection is positive, this validates the purity of the subset Si to be tested as reliably (as far as purity validation is concerned) as if the test had been carried out on the basis of the complete basic temperature fields.
[0186] Indeed, a mathematical demonstration shows that: • if the convex hull Ëi of the set grouping the temperature subfields Tp that are the projections of the basic temperature fields T; belonging to a subset Si, contains no temperature subfield that is associated with a basic temperature field Tr of a cluster other than that of the subset Si, • then, the convex hull Ei of the subset Si contains no base temperature fields T; • of a cluster other than that of the subset Sp
[0187] That is to say, in summary: if the projection of the subset Si is pure, then the subset Si itself is also pure.
[0188] The demonstration of this property is not included here for the sake of brevity, and because it is not of direct interest for the implementation of the invention.
[0189] It should be noted, however, that the converse of this proposition is generally not true. Indeed, if the projection of the subset Si is not pure, one cannot generally deduce anything from it concerning the purity or impurity of the subset Si itself.
[0190] Thus, the projection of the subset Si may very well not be pure, even though the subset itself is pure. Indeed, it is clear that a projection, which constitutes a kind of compression along certain directions, can very well bring into a subset elements belonging to another cluster, which, before the projection, were outside that subset. In other words, a projection can very well cause a loss of purity.
[0191] To take into account these two aspects (namely: the purity of the projected implies the purity of the subset Si itself; and the non-purity of the projected implies uncertainty as to the purity of the subset Si), the purity test step Al 11 is carried out as follows ([Fig. 13]): • we execute a test step by projection Al 12; • when the result of the Al 12 projection test is positive (i.e., when the projected subset Si is pure), we determine that the subset Si is pure; • On the other hand, when the result of the Al 12 projection test is negative, we perform this projection test, but by carrying out the projection on the basis of another subspace (of another restricted group of nodes, selected randomly).
[0192] As long as the result of the projection test Al 12 is negative, and as long as the number it of iterations of this test is less than a limit number of iterations, niim, the step Al 12 is executed again.
[0193] If the result of the Al 12 projection test eventually becomes positive, during these iterations, it is determined that the subset Si is pure.
[0194] On the other hand, if after niim iterations of the projection test (each time projecting onto a different subset), the test still does not yield a positive result, it is determined that the subset Si to be tested is not pure. In reality, it is unknown whether it is pure or not. However, since it could not be ascertained that it is pure, this subset is eliminated as a precaution, and no additional temperature field will be determined within it.
[0195] In practice, the Al 12 projection test can be carried out as explained now.
[0196] During the Al 12 projection test, it is determined, for the subset Si whose purity is to be tested, whether the convex envelope Ëj of the set grouping the temperature subfields Tj, which are the projections of the basic temperature fields T; of the subset Si, does not contain any temperature subfield Tj, which is associated with a foreign basic temperature field Tr, of a different cluster Ck- than the cluster Ck to which the subset Si to be tested belongs.
[0197] This amounts to testing that no temperature subfield Tj can be written as a linear combination of the temperature subfields Tj of the subset Si, assigned positive coefficients and normalized to 1. To test this criterion, we proceed as follows.
[0198] For each temperature subfield Tj, corresponding to one of the basic temperature fields Ti of the subset Si, a vector jq is considered having QR+1 components. The first QR components of the vector jq are the QR temperature values given by the temperature subfield Tj. The last component of this vector is equal to 1.
[0199] We then form a matrix A, of dimension (QR+l)xn, by grouping in matrix form the n column vectors jq corresponding to the n temperature subfields T} that comprise the subset Si whose purity we want to test.
[0200] For each temperature subfield Tj belonging to a different cluster Ck- than that of the subset Si to be tested, a vector jti- is constructed in the same way.
[0201] To test whether the temperature subfield Tj is indeed different from any linear combination of the temperature subfields Tj of the subset Si to be tested, assigned positive coefficients and normalized to 1, the following condition is then tested:
[0202] min || A w - nj, || > £ || Hj, || w^R"
[0203] where w is a vector grouping n positive coefficients and where e is a fixed tolerance level, for example between 0.5% and 5%.
[0204] From a numerical point of view, to test this condition, a least-squares minimization (with positive coefficients) is performed on the left-hand side. The coefficients grouped in the vector w are not normalized to one (removing this constraint simplifies the minimization procedure in question). However, the component equal to 1, added to the end of each vector ir, allows this test to be performed as if testing a linear combination with positive coefficients normalized to 1.
[0205] Numerical example
[0206] Results obtained using this method of determining training data are presented briefly below as an example of implementation.
[0207] This example corresponds to the part shown in [Fig. 5], for which the high-resolution mesh comprises QH = 42,445 nodes. Furthermore, N = 600 high-resolution temperature fields TH.i, ..., TH, n are initially produced in step 10. And the number P of clusters (i.e., the number P of response types Mk) is equal to 4 here.
[0208] As already indicated, in this example, the variable selection step S0 allows, in the high-resolution mesh, the selection of Q=87 relevant nodes, which allows the high-resolution temperature fields to be simplified into the basic temperature fields Tb..., TN, still numbering 600.
[0209] Next, during the data augmentation step A0, approximately 60 pure subsets were identified (on average) in each cluster Ci,...CP of temperature fields. Each pure subset thus identified comprises on average 5 elements (5 basic temperature fields). During the purity test, by projection (step Al 12), the subspaces onto which the temperature fields are projected are 5-dimensional subspaces (in other words, 5 nodes, instead of 87, for the purity test by projection).
[0210] In this example, 5400 additional temperature fields T';, each with its label, were then obtained by convex linear combination within these pure subsets.
[0211] The set of training data obtained (basic + additional) then includes 6000 labeled temperature fields, used for training a temperature classifier.
[0212] The classification accuracy obtained after training was then tested for several types of classifiers: • with and without data augmentation (with / without DA), and • with a selection of variables: • either by the selection technique described above, by maximizing relevance while minimizing redundancy (FS), • either by principal component analysis (PCA).
[0213] These results are summarized in the table below. Classifier Type Used Dimension Reduction Accuracy, With DA Accuracy, Without DA Multilayer Perceptron FS 87.0% 81.0% PCA 86.5% 81.5% Quadratic Discriminant Analysis FS 77.5% 70.5% PCA 76.0% 70.0% Gaussian Naive Bayes FS 39.5% 34.5% PCA 38.5% 31.5%
[0214] These results show that additional training data, as determined here, do indeed improve classification accuracy, and are therefore physically relevant.
[0215] These results also show that the variable selection (FS) method presented above is reliable (in addition to being fast to execute), in that it effectively allows the most relevant information to be extracted from high-resolution temperature fields.
[0216] Electronic devices
[0217] The present technology also relates to an electronic device comprising at least one processor and one memory, programmed to execute the variable selection method and / or the data augmentation method described above. More generally, this electronic device can be programmed to execute the entire method for determining training data and training the classifier of [Fig. 4].
[0218] The device in question may be implemented in the form of a computer or a computing unit whose components are located in the same place, for example in the same room or in the same computer case. It may also be implemented in the form of a decentralized computing system that uses, via a computer network, remote, distributed computing resources (for example, remote computing resources provided by a "cloud").
[0219] The present technology also relates to an electronic device, comprising at least one processor and one memory, comprising a classifier, i.e., programmed to perform a classification function as described above, and in which the classifier has been trained on the basis of the basic and / or supplementary training data presented above. This classifier is specific (and, as seen above, particularly accurate) due to the specific training it has undergone. From a practical point of view, the specific characteristics of this classifier are reflected in the particular values of the coefficients (synaptic coefficients, for example) which parameterize the classifier, as obtained after training on the basis of the basic and / or additional training data.
[0220] The present technology also relates to an electronic device, comprising at least a processor and a memory, programmed to implement a numerical simulation method of the mechanical response of a turbine blade element subjected to a given temperature field TR, the simulation method comprising: • a preliminary training phase including: • a determination of baseline training data and supplementary training data, as explained above, • training a temperature field classifier based on said basic training data and said additional training data, and • a usage phase comprising the following steps: • identification of one of the standard responses Mb ..., MP of said blade element by the temperature field classifier receiving as input said temperature field TR, and • numerical simulation of the mechanical response of said blade element, subjected to said temperature field TR, in accordance with the simplified mechanical behavior model defined by the standard response previously identified by the classifier.
Claims
1. Demands A computer-implemented method for the numerical simulation of the mechanical response of a turbine blade element subjected to a given stress field (TR), comprising: - a preliminary training phase comprising: • a determination of baseline training data for training a stress field classifier (2'): • wherein high-resolution training data gather high-resolution stress fields (TH>b TH>i, TH N) exerted on a turbine blade element (1' ; 1) and, for each of these fields, a label (k) associated with the stress field considered, each high-resolution stress field (TH,i) giving values of said stress (T'hj, TjH>i, Tj H>i) at different nodes of a high-resolution mesh of said blade element, the label (k) associated with the high-resolution stress field (TH>i) considered identifying a typical response (Mk), representative of an expected mechanical response for said blade element (1' ; 1) in response to said high-resolution stress field (TH>i), among different typical responses (Mb Mk, MP) of said turbine blade element, • the method comprising a selection step (SO) of a limited number (Q) of nodes from the high-resolution mesh, the selected nodes forming a group of relevant nodes (S),
2. • the basic training data, which includes basic constraint fields (Tb, T), TN) obtained by restricting the high-resolution constraint fields (TH,b, TH.i, THjn) to the relevant node group, • and in which at least some of said relevant nodes are selected each on the basis of an approximate value (7(dJJ')) of a mutual redundancy information (Ijj) between the node (j) considered and any of the other nodes (j') of said mesh, said approximate value being determined on the basis of a spatial distance (djj) between the two nodes considered, • the training (20) of a constraint field classifier on the basis of said basic training data, or on the basis of said basic training data and additional training data, - a usage phase comprising the following steps • identification of one of the standard responses (Mb Mk, MP) of said blade element by the stress field classifier receiving said stress field (TR) as input, • numerical simulation of the mechanical response of said blade element, subjected to said stress field (TR), in accordance with the simplified mechanical behavior model defined by the standard response previously identified by the classifier. Method according to the preceding claim, wherein said stress exerted on the turbine blade element (1'; 1) is a temperature exerted on this element, the high-resolution stress fields (TH,b TH.i, THjn) and the basic stress fields (Ti, Tn, Tn) being respectively high-resolution temperature fields and basic temperature fields.
3. A method according to any one of the preceding claims, wherein the mutual redundancy information (Ijj) between the two nodes (j, j') considered is representative of a level of information redundancy between: - the values of said constraint (TjH> b TjH>i, TjH> N) at the first (j) of the two nodes, as given by the set of high-resolution constraint fields (Th.i, TH,i, Th>n), and - the values of said constraint (Tj H, i, Tj Hi, Tj H, n) at the second (j') of the two nodes, as given by the set of high-resolution constraint fields (Th,i, TH>i, THjN).
4. A method according to any one of the preceding claims, wherein the preliminary training phase includes the determination of The basic training data includes a calibration step (SU) which comprises the following steps: - (S12) for each pair of nodes (j, j') of a subset grouping a portion of the nodes of the high-resolution mesh: • determination of a reference value (j[) of the mutual redundancy information (Ijj) between the two nodes (j, j') considered, by statistical calculation based on the values of said constraint (TjH>i, TjH>N, TjH, i, TjH>i, TjH>N), at the level of these two nodes, as given by the different high-resolution constraint fields (TH4, TH>i, THN), and • calculation of the spatial distance (djj) separating the two nodes (j, j') considered - (S13) fitting of a curve (i), so that it is representative of a link between said reference values (j(pd pJ J)) of the mutual information of redundancy, and the corresponding spatial distances (djj),- and in which the approximate value (f, J) of the mutual redundancy information between two nodes (i, i') of the network is the value given by said curve (i), for the spatial distance (djj) separating the two nodes considered.
5. A method according to any one of the preceding claims, wherein each relevant node (j) is further selected on the basis of a degree of relevance ( ) of the value of said stress at the node (j) considered, with respect to the mechanical behavior of said blade element, said degree of relevance ( ) being representative of a level of dependence between, on the one hand, the label (k) associated with the high-resolution stress field (TH,i) considered, and, on the other hand, the value of said stress (TjH>i) at the node (j) considered for this stress field.
6. Method according to the preceding claim, wherein: - the relevant node group (S) initially contains a first node, which is the node (j) of the high-resolution mesh having the highest degree of relevance ( - an addition step is then performed, during which an additional node is added to the relevant node group (S), the additional node being selected, from among the nodes of said mesh which are not yet part of said group, so as to maximize an amount equal to: the degree of relevance ( f ( ) of the node considered, less an average level of redundancy between this node and the nodes already present in the relevant node group (S), - the addition step is then performed again several times.
7. Method according to the preceding claim, wherein the average level of redundancy, between the node under consideration and the nodes already present in the relevant node group (S), is equal to the arithmetic mean of the approximate values (I) of the mutual redundancy information (hj) between said node (i) and the nodes (j) already present in the relevant node group (S), the additional node, added to the relevant node group (S), being the node that maximizes the following quantity: T ( 1 X ÏM 1 where i is the index locating 1P^1 > ~ Card(S) ) the node under consideration, in the high-resolution mesh, and where Cârd(S) is the number of nodes contained in the relevant node group (S) during the current iteration of the addition step.
8. A method according to any one of the preceding claims, wherein the preliminary training phase comprising the determination of basic training data further comprises a data augmentation step (AO) during which the additional training data are determined from said basic training data, each basic constraint field (Tb T, TN) being associated with a label (k) of the field
9. corresponding high-resolution constraint (Th,i, TH>i, THN), the data augmentation step includes the following steps: • (Al) form at least one pure subset (Si), which groups basic stress fields (T,) associated with the same standard response (Mk) of the turbine blade element, and which has a convex envelope (Ei) containing no basic stress field (T; ) that is associated with another standard response (Mk ) of said element, • (A2) for at least one of said pure subsets (Si), determination of the additional training data, each additional training data comprising an additional constraint field (T'i), obtained by linear combination of basic constraint fields (T,) of the pure subset (Si) considered, assigned positive weighting coefficients and normalized to 1, the additional constraint field (T'i) being accompanied by a label (k) which is the same as for the basic constraint fields (T,) of said combination. An electronic device comprising at least one processor and one memory, programmed to execute the numerical simulation method according to one of the preceding claims.