A key area identification method based on a four-dimensional characterization index

By constructing a four-dimensional characterization index system and combining Latin hypercube sampling and Pareto dominance relations, the problems of accuracy and efficiency in identifying key areas in digital twin experiments of aero-engines were solved, and the rational allocation of experimental resources and cost reduction were achieved.

CN122113449AActive Publication Date: 2026-05-29NAT UNIV OF DEFENSE TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NAT UNIV OF DEFENSE TECH
Filing Date
2026-04-21
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing key area identification methods lack unified measurement standards in digital twin tests of new aero-engines, making it impossible to accurately capture gradient abrupt changes near the compressor surge line and local extreme values ​​of turbine blade thermal load. Furthermore, the identification results of different methods are difficult to compare horizontally, leading to unreasonable allocation of test resources and low resource utilization efficiency.

Method used

Sample points are generated using Latin hypercube sampling. Response values ​​are calculated using a digital space-driven engine performance model. A four-dimensional representation index vector is constructed, which includes information surface entropy, maximum gradient consistency, local-global similarity divergence, and the ratio of local to global extreme points. Key regions are selected by combining Pareto dominance and non-dominated sorting.

Benefits of technology

It achieves comprehensive capture of key regional features, provides unified quantitative standards and verification basis, guides the precise allocation of experimental resources, improves experimental efficiency and reduces costs, and enhances the utilization efficiency of digital twin experimental resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113449A_ABST
    Figure CN122113449A_ABST
Patent Text Reader

Abstract

The application relates to a key region identification method based on a four-dimensional characterization index. The method comprises the following steps: based on a sample set, calculating an information surface entropy, a maximum gradient consistency, a local-global similarity divergence and a local-global extreme point ratio to form a four-dimensional characterization index vector of a key region; fusing the four-dimensional characterization index vector based on a Pareto dominance relationship, dividing candidate key regions into different front levels through non-dominated sorting, calculating the crowding distance of each front region, and screening the key region in the test sample space in combination with the front level and the crowding distance. The method can improve the utilization efficiency of a new type of aero-engine digital twin test resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital equipment testing technology, and in particular to a method for identifying key areas based on four-dimensional characterization indicators. Background Technology

[0002] In the digital twin testing of new aero-engines, it is necessary to drive the engine performance model through digital space, and combine digital testing with semi-physical simulation to complete the assessment of compressor surge boundaries, turbine blade thermal load limits, and overall engine performance margin. Aero-engine digital equipment testing involves multiple high-dimensional parameters such as compressor speed, inlet temperature, pressure ratio, fuel flow rate, turbine inlet temperature, and cooling gas flow rate, forming a test sample space with dozens of dimensions. Due to the dual limitations of test bench resources, test cycle, and computational cost, it is impossible to fully traverse the high-dimensional test sample space. Therefore, it is necessary to identify the key areas that play a decisive role in the confidence of the engine performance model, and concentrate limited test resources on key areas such as surge boundaries and thermal load critical points to achieve model confidence verification requirements at low cost.

[0003] Key region identification is a core supporting link in the entire lifecycle verification and validation of digital equipment models for aero-engines, requiring a complete process of test sample space analysis, response function exploration, key region identification, and identification effect verification. However, existing key region identification methods have core technical defects: they lack a unified measurement standard for the amount of information in key regions, relying solely on expert experience to judge factor boundaries on compressor characteristic maps or simply screening high-fluctuation areas through sample variance, failing to comprehensively and quantitatively characterize the performance surface ruggedness, gradient change characteristics, and local-global distribution differences of key regions; simultaneously, the surrogate model for aero-engine digital equipment testing is a black-box function without analytical expression support, making it difficult for existing methods to accurately capture key features such as gradient abrupt changes near compressor surge lines and local extrema of turbine blade thermal load, and the lack of unified verification criteria for key regions identified by different methods makes it impossible to scientifically judge their effectiveness and compare results horizontally, leading to unreasonable allocation of test resources, low accuracy and efficiency in key region identification, and low utilization efficiency of resources for new aero-engine digital twin testing. Summary of the Invention

[0004] Therefore, it is necessary to provide a key area identification method based on the construction of four-dimensional characterization indicators to address the above-mentioned technical problems and improve the utilization efficiency of digital twin test resources for new aero-engines.

[0005] A key region identification method based on a four-dimensional representation index, the method comprising:

[0006] In the test sample space of digital equipment testing, sample points are generated by Latin hypercube sampling. Each sample point is calculated to obtain the corresponding response value through the engine performance model driven by digital space. The response value is at least one of the aero-engine performance indicators. A sample set containing sample points and corresponding response values ​​is constructed. Each dimension of the test sample space is composed of at least one of the aero-engine physical parameters. Based on the sample set, the information surface entropy, maximum gradient consistency, local-global similarity divergence and local-global extreme point ratio are calculated to form a four-dimensional characterization index vector of the key region. Based on the Pareto dominance relation, the four-dimensional characterization index vector is fused. The candidate key regions are divided into different frontier levels through non-dominated sorting. The crowding distance within each frontier region is calculated. The key regions in the test sample space are obtained by combining the frontier level and the crowding distance. The key regions correspond to the mapping subspace of surge boundary region, thermal load critical region, performance inflection point region or stability transition region in the test sample space.

[0007] The aforementioned key region identification method based on four-dimensional characterization indices, this application constructs a four-dimensional characterization index system including ILEM, MGCM, LGSD, and LGER. It comprehensively quantifies the information characteristics of key regions from four orthogonal dimensions: ruggedness, rate of change, distribution difference, and complexity. This overcomes the limitations of traditional single-indicator or experience-based judgments, achieving comprehensive capture of key region characteristics. Each index is based on Latin hypercube sampling to ensure statistical representativeness of the calculation, adapting to the non-analytical nature of the black-box function in digital equipment testing. Furthermore, a rigorous mathematical calculation model is designed to ensure the objectivity, calculability, and repeatability of the index values. Pareto dominance relationships enable the fusion of indicators without subjective weights. By combining non-dominated ranking and crowding distance, key regions are objectively screened, providing a unified quantitative standard and verification basis for key region identification. This solves the problem of difficulty in comparing the results of different identification methods. The first frontier high-crowding distance region selected is highly matched with the core requirements of the boundary performance to be tested in digital equipment tests. This can guide the precise allocation of test resources, significantly improve test efficiency, reduce test costs, and provide core technical support for the confidence verification of the digital equipment model of the new aero-engine. This also improves the utilization efficiency of digital twin test resources for the new aero-engine. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating a key region identification method based on a four-dimensional characterization index in one embodiment. Detailed Implementation

[0009] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0010] In one embodiment, such as Figure 1 As shown, a key region identification method based on a four-dimensional representation index is provided, including the following steps: Step 102: In the test sample space of digital equipment testing, sample points are generated by Latin hypercube sampling. Each sample point is calculated to obtain the corresponding response value through the engine performance model driven by digital space. The response value is at least one of the aero-engine performance indicators. A sample set containing sample points and corresponding response values ​​is constructed. Each dimension of the test sample space is composed of at least one of the aero-engine physical parameters.

[0011] The test sample space consists of high-dimensional parameters affecting engine performance. Each dimension of the test sample space is composed of at least one of the physical parameters of the aero-engine, and each sample point corresponds to an engine operating point. Specifically, the sample points include parameters such as compressor high-pressure rotor speed, low-pressure rotor speed, compressor inlet total temperature, turbine inlet total temperature, compressor pressure ratio, fuel flow rate, turbine blade cooling air flow rate, adjustable stator blade angle, flight altitude, and flight Mach number. Each sample point is simulated and calculated using a digital space-driven engine performance model to obtain corresponding response values, including performance indicators such as engine thrust, fuel consumption rate, compressor surge margin, maximum turbine blade surface temperature, compressor efficiency, and turbine efficiency. Latin hypercube sampling is employed to generate sample points in a high-dimensional parameter space, ensuring uniform coverage of samples across each dimension. The constructed sample set consists of several sets of sample points and their corresponding response values. Each set of sample points represents a specific engine operating state, providing a data foundation for the subsequent calculation of four-dimensional characterization indicators. This supports the identification of key regions such as surge boundary regions, thermal load critical regions, performance inflection point regions, and stability transition regions. During the calculation process, the experimental sample space is... d In high-dimensional spaces, Latin hypercube sampling (LHS) ensures that sample points are uniformly distributed in the high-dimensional space, exhibiting good spatial filling and statistical representativeness; each sample point contains d The coordinate values ​​of each dimension are used to calculate the response value corresponding to each sample point through the experimental objective function or black box response function, forming a sample set consisting of the coordinates of the sample points and their corresponding response values. Each row of the sample set represents a sample point and its response value, providing basic data support for the subsequent calculation of four-dimensional characterization indicators, thereby supporting the identification of key regions such as surge boundary region, thermal load critical region, performance inflection point region and stability transition region.

[0012] Specifically, in the experimental sample space, there is a sample set. , For the objective function or response function,

[0013] in, Each row represents a sample point. d Represents the dimension of space. n This represents the sample size. Each sample point... It is achieved through Latin hypercube sampling (LHS) in the region Internally generated.

[0014] Step 104: Based on the sample set, calculate the information surface entropy, maximum gradient consistency, local-global similarity divergence, and local-global extreme point ratio to form a four-dimensional representation index vector for the key region.

[0015] The four-dimensional characterization index vector contains four aspects of feature description: Information Surface Entropy (ILEM), Maximum Gradient Consistency (MGCM), Local-Global Similarity Divergence (LGSD), and Local-Global Extremum Ratio (LGER). Each sub-index reflects the local and global features in the experimental sample space, respectively.

[0016] In digital equipment testing of aero-engines, the area near the compressor surge boundary often exhibits characteristics such as drastic fluctuations in the response surface (high information surface entropy), gradient mutations (high maximum gradient consistency), significant differences between local and global operating condition distributions (high local-global similarity divergence), and dense extreme points (high local-global extreme point ratio). The above four indicators quantify the key region characteristics from four orthogonal dimensions: ruggedness, rate of change, distribution difference, and complexity. They are combined to form a four-dimensional characterization index vector, which enables a comprehensive and objective characterization of the information content of key regions such as surge boundary and thermal load critical point, and is adapted to the characteristics of the black box function without analytical expression in digital equipment testing.

[0017] Step 106: Based on the Pareto dominance relation, the four-dimensional characterization index vector is fused. The candidate key regions are divided into different frontier levels by non-dominated sorting. The crowding distance of each frontier region is calculated. The key regions in the test sample space are obtained by combining the frontier level and the crowding distance. The key regions correspond to the mapping subspace of the surge boundary region, thermal load critical region, performance inflection point region or stability transition region in the test sample space.

[0018] In digital equipment testing of aero-engines, tasks such as surge boundary assessment and turbine blade thermal load limit testing require prioritizing the allocation of limited test bench and computing resources to critical areas that have the greatest impact on model confidence. This involves quantifying the index vector using the concept of non-dominated solutions and processing the solution set. Any two key regions and The Pareto dominance relationship is used to determine the quality of the solutions. When regions belong to different frontiers, the solution set is divided into multiple frontiers according to the dominance relationship. The first frontier contains all regions not dominated by other solutions, the second frontier contains regions dominated only by solutions from the first frontier, and so on. A region located at a higher frontier indicates that its overall performance is better. Crowd distance measures the relative position of the critical region on various indicators. A larger crowd distance value indicates that the region is located at the edge or in a sparse area of ​​the indicator space, corresponding to the vicinity of the performance boundaries of aero-engines (surge boundary, thermal load limit boundary, etc.). These regions are the key areas that need to be verified in digital equipment testing. By combining Pareto dominance, non-dominated ranking, and crowd distance, critical regions can be effectively screened and evaluated. This step selects regions in the first frontier as critical regions to guide subsequent test resources to be tilted towards high-risk, high-value regions such as surge boundary regions and thermal load critical regions.

[0019] Specifically, the index vector is quantified by drawing on the idea of ​​non-dominated solutions, and the solution set is... Any two key regions and Through Pareto dominance Judging its merits and demerits: .

[0020] In a set of sub-regions within the experimental sample space, with only two solutions... and In this case, the quantitative comparison of key regional indicators can be simplified into the following three scenarios: First, and In this case, Considered superior .

[0021] second, and In this case, Considered superior .

[0022] Third, there is at least one indicator. k Make And there is another indicator. l Make In this case, and They belong to different frontiers, and it is impossible to directly compare their advantages and disadvantages.

[0023] When the regions belong to different frontiers, the solution set will be... Divided into multiple frontiers according to dominance relationship ,in The first frontier comprises all regions not dominated by other solutions; The second frontier contains regions dominated only by solutions from the first frontier, and so on.

[0024]

[0025]

[0026] A region located at a higher frontier indicates that its overall indicators are better. For each frontier... The solution is used to calculate its crowding distance in the index space. :

[0027] The crowding distance of the boundary solution is set to infinity:

[0028] Crowded distance Key areas were measured Relative position on various indicators. Larger Values ​​representing regions located at the edges or in sparse areas of the indicator space may represent performance boundaries of interest in experimental design. Combining Pareto dominance, non-dominated ordination, and crowding distance allows for effective screening and evaluation of critical regions. This step involves selecting... The area in the middle is the key area.

[0029] The aforementioned key region identification method based on four-dimensional characterization indicators constructs a four-dimensional characterization indicator system including ILEM, MGCM, LGSD, and LGER. This system comprehensively quantifies the information characteristics of key regions from four orthogonal dimensions: ruggedness, rate of change, distribution difference, and complexity. This overcomes the limitations of traditional single indicators or empirical judgments, achieving comprehensive capture of key region characteristics. Each indicator is based on Latin hypercube sampling to ensure statistical representativeness, adapting to the non-analytical nature of the black-box function in digital equipment testing. A rigorous mathematical calculation model is designed to ensure the objectivity, calculability, and repeatability of indicator values. Based on Pareto dominance relations, subjective weightless fusion of indicators is achieved. Combined with non-dominated ranking and crowding distance, objective screening of key regions is completed, providing a unified quantitative standard and verification basis for key region identification and solving the problem of difficult horizontal comparison of identification results from different methods. The selected first-frontal high-crowding-distance region highly matches the core requirements of the key boundary performance verification in digital equipment testing, guiding the precise allocation of test resources, significantly improving test efficiency, reducing test costs, providing core technical support for the confidence verification of digital equipment models, and improving the utilization efficiency of digital twin test resources for new aero-engines.

[0030] In one embodiment, the calculation process of the information surface entropy includes: A walk sequence is generated based on space-filling design, and the optimal point in the sample set is determined based on the walk sequence. ,in, Describe the objective function. Representing the sample space, Indicates sample points; Construct the adaptation vector and reference vector based on the optimal point; The entropy of the information surface is calculated based on the adaptation vector and the reference vector:

[0031] in, The number of sample points in the walk sequence. To adapt the first vector i The component and the first j The measurement between the components For the reference vector, the first i The component and the first j The measurement between the components.

[0032] Specifically, , ,

[0033] in, Represents sample points The objective function value, Represents sample points The objective function value, Indicates the tolerance threshold. Indicates the best point A quadratic reference function centered at the center is used to construct the reference vector; Information surface metric is a method for evaluating the difficulty of solving a problem by an algorithm based on the information distribution in the solution space. Based on the theory of information surface metric, this application proposes information surface entropy, using a walk sequence generated based on space-filling design as the basis for calculation, ensuring efficient coverage and representativeness of the sample distribution. In the minimization problem, the optimal point in the optimal point sampling set is first determined, and then an adaptation vector and a reference vector are constructed. Information surface entropy can reflect the diversity and uncertainty of a region. A larger information surface entropy value indicates a higher degree of ruggedness in the region; conversely, a smaller information surface entropy indicates a smoother space and lower ruggedness. This index can effectively quantify the ruggedness and information dispersion of the regional response surface, providing a multi-dimensional characterization basis for key region identification. In aero-engine digital equipment testing, a larger information surface entropy indicates a more rugged compressor surge margin or turbine blade temperature response surface in that region, and a greater likelihood of containing surge boundaries or thermal load critical points.

[0034] In one embodiment, the calculation process for maximum gradient consistency includes: Based on the walking sequence generated by Latin hypercube sampling, a normalized scale is defined to eliminate the influence of dimensions, and a minimum value is set to avoid the division by zero problem; The mean of the dot product of the gradient vectors of the sample points in the sample set is calculated based on the normalized scale, yielding the maximum gradient consistency as follows:

[0035] in, For the first i The difference in response values ​​between the groups of samples. For sample points, For normalization, For the maximum response value, For the minimum response value, For the set minimum value, Step size, This represents the number of sample points in the walk sequence.

[0036]

[0037] in, Indicates the first sample in the sample set k The sample point at the th th j Coordinates in each dimension. To avoid division by zero, a minimum value is set. .

[0038] Specifically, in the experimental sample space, maximum gradient consistency is defined as a key indicator for judging the local rate of change in a region. The walk sequence generated by Latin hypercube sampling is used as the basis for calculating maximum gradient consistency, thus ensuring the uniformity and representativeness of the sample distribution in the high-dimensional search space. The normalization scale is defined to eliminate the influence of dimensions, while a minimum value is set to avoid the division-by-zero problem. The smaller the value of maximum gradient consistency, the more gradual the change within the region; conversely, the larger the value of maximum gradient consistency, the more drastic the change within the region, and the more likely this region is a critical area. This indicator can effectively characterize the drastic degree of local rate of change of the response function within a region. This indicator is used to capture the locations of gradient abrupt changes on the compressor characteristic diagram, such as sharp changes in thrust or pressure ratio near the surge line. Regions with higher maximum gradient consistency are more worthy of focused experiments.

[0039] In one embodiment, the calculation process of the local-global similarity divergence includes: The search space is divided into equally probable intervals using Latin hypercube sampling, and sample points are generated in both the global space and local regions, including: Through LHS d The search space is divided into several dimensions. There are three equally probable intervals, and points are randomly selected within each interval to ensure uniform coverage of the samples in each dimension. Pre-generated in the original search space... High-dimensional sample points At the same time, the current search space is set as:

[0040] Then through LHS Intrinsic generation A high-dimensional sample point.

[0041] For each generated sample point Its corresponding volume element Determined by the following formula:

[0042] in, Let be the total volume of the search space. Due to the uniform partitioning of the LHS along each dimension, all volume elements... They are mathematically equal; Gaussian kernel density estimation is used to estimate the probability density functions of the local distribution and the global reference distribution, respectively; The local-to-global similarity divergence is calculated based on KL divergence as follows:

[0043] in, For local distribution at grid points The probability density estimate at that location, This is an estimate of the probability density of the global reference distribution at the same grid point. For the first k Volume element of each grid point The total number of grid points. These are discrete grid points.

[0044] Specifically, to characterize the similarity between local and global regions in the experimental sample space, a local-global similarity divergence index is proposed based on the theory of KL divergence. The search space is divided into d-dimensional equal-probability intervals using Latin hypercube sampling, and points are randomly selected within each interval to ensure uniform coverage of samples in each dimension. High-dimensional sample points are pre-generated in the original search space, and the current search space is defined. Then, high-dimensional sample points are generated within the current search space using Latin hypercube sampling. The probability density functions of the local and global reference distributions are estimated using Gaussian kernel density estimation, and then the local-global similarity divergence is calculated based on KL divergence. The larger the local-global similarity divergence, the greater the difference between the data distribution within the region and the overall spatial distribution, and the more likely it is to be a key region of interest.

[0045] When the sample distribution in a local operating condition area differs significantly from the spatial distribution in the entire operating condition area, it indicates that the area has special performance characteristics, such as surge boundary offset at high altitude and low Mach number, which needs to be verified separately.

[0046] In one embodiment, Gaussian kernel density estimation is used to estimate the probability density functions of the local distribution and the global reference distribution, respectively, including: The probability density estimate of the local distribution is calculated as follows:

[0047] The probability density estimate of the global reference distribution is calculated as follows:

[0048] in, and For different bandwidth parameters, For Gaussian kernel function, For spatial dimensions, For local area sample points, For global spatial sample points, m This represents the total number of sample points in the global space. j The index number of the sample point.

[0049] Specifically, for each generated sample point, its corresponding volume element is determined by dividing the total volume of the search space by the total number of grid points. Since Latin hypercube sampling provides a uniform partition in each dimension, all volume elements are mathematically equal. The bandwidth parameter is selected based on the regularized bandwidth selection method proposed by Silverman, and the Gaussian kernel function takes the following form: ,in, The standard deviation of the sample. n The sample size is given. Gaussian kernel density estimation allows for the estimation of the probability density functions of both the local and global reference distributions, providing a foundation for subsequent KL divergence calculations.

[0050] In one embodiment, the calculation process for the local-to-global extreme point ratio includes: The BFGS algorithm is used to search for local extrema within a local region, and the number of local extrema is counted, including: For local extrema, the BFGS algorithm is used to find and count them. The number of local extrema is... The formal representation is:

[0051] For each initial point The extreme points it searched for were ,symbol Indicates the number of elements; Count the number of local extrema within the global scope, and calculate the ratio of local to global extrema:

[0052] in, Indicates a local area The number of local extrema within, Indicates global scope The number of local extrema within, That is, the local region is a subset of the global region.

[0053] Specifically, the number of local extrema is often used to quantify the complexity of the function in the unset region. For local extrema, the BFGS algorithm is used to find and count them. The formal representation of the number of local extrema is the counting of the extrema found for each initial point. By calculating the ratio of the number of extrema in the local region to the global region, a local-global extrema ratio index is proposed to measure the distribution characteristics of the function and the importance of local features. When the local and global regions use the same sampling scale, and the local-global extrema ratio approaches 1, it indicates that the density of extrema in the local region is higher than the global average. This corresponds to a dense area of ​​extreme heat load on turbine blades or an inflection point region of compressor efficiency, which is a key verification object in digital equipment testing.

[0054] In one embodiment, the four-dimensional representation index vector is fused based on the Pareto dominance relation, including: For any two candidate key regions, compare their four-dimensional representation index vectors. If all indices of one region are not inferior to those of the other region and at least one index is superior to that of the other region, then the former is determined to dominate the latter.

[0055] Specifically, in the set of sub-regions of the experimental sample space, when there are only two solutions, the quantitative comparison of key region indicators can be simplified to the following three cases: First, all indicators of one region are better than those of another region, in which case the former is considered better than the latter; second, all indicators of one region are worse than those of another region, in which case the latter is considered better than the former; third, there exists at least one indicator that makes the former better than the latter, and there exists another indicator that makes the former worse than the latter, in which case the two regions belong to different frontiers and cannot be directly compared in terms of their superiority or inferiority. Through Pareto dominance, subjective weightless fusion of indicators can be achieved, providing a basis for the objective selection of key regions.

[0056] In one embodiment, candidate key regions are divided into different frontier levels through non-dominated sorting, including: Regions not dominated by any other region are designated as the first frontier; regions dominated only by the first frontier region are designated as the second frontier, and so on, until all candidate key regions are assigned to the corresponding frontier level.

[0057] Specifically, when regions belong to different frontiers, the solution set is divided into multiple frontiers according to dominance relationships. The first frontier contains all regions not dominated by other solutions, the second frontier contains regions dominated only by solutions from the first frontier, and so on. A region located at a higher frontier indicates that its overall performance is better. Through non-dominated ranking, candidate key regions can be automatically divided into different frontier levels, providing a hierarchical basis for subsequent key region selection.

[0058] In one embodiment, the crowding distance within each frontal region is calculated, and key regions in the experimental sample space are selected by combining the frontal level and the crowding distance, including: For each region within the same frontier, its crowding distance in the index space is calculated as follows:

[0059] in, For the region Crowded distance, For the first j Under these indicators, the region The value of , For the first j The maximum value of each indicator. For the first j The minimum value of each indicator. For the j-th indicator, the region The value of , To exclude the frontier The two boundary areas at the beginning and end; By combining the frontier hierarchy and crowding distance, key regions in the experimental sample space were identified, including: Prioritize regions within the first frontier as candidate key regions; Within the first frontier, areas are sorted from largest to smallest based on their crowding distance, and the areas with the largest crowding distance are selected as the final key areas.

[0060] Specifically, for each solution in the frontier, its crowding distance in the indicator space is calculated. The crowding distance for boundary solutions is set to infinity. The crowding distance measures the relative position of the critical region on various indicators. A larger crowding distance value indicates that the region is located at the edge or in a sparse region of the indicator space, which may be the performance boundary of interest in the experimental design. By combining the frontier hierarchy with the crowding distance, critical regions can be effectively screened and evaluated, and regions in the first frontier can be selected as critical regions. The four-dimensional indicators provide multi-dimensional features of regional information because information surface entropy, maximum gradient consistency, local-global similarity divergence, and local-global extreme point ratio characterize regional characteristics from four orthogonal dimensions: ruggedness, rate of change, distribution difference, and complexity, respectively. Therefore, different types of critical regions can be accurately captured. Compared with a single variance indicator, the sensitivity for identifying nonlinear abrupt change regions is improved; the Pareto frontier hierarchy enables automatic sorting of regions, allowing for intuitive selection of critical regions.

[0061] In digital equipment testing of aero-engines, surge boundary regions may simultaneously exhibit high information surface entropy (rugged surface) and high maximum gradient consistency (gradient abrupt change), while thermal load critical regions may simultaneously exhibit high local-global similarity divergence (abnormal distribution) and high local-global extreme point ratio (dense extreme points). Different critical regions exhibit varying degrees of emphasis on different indicators, making it difficult to objectively determine the weight of each indicator using traditional weighted summation methods. This application employs Pareto dominance to achieve indicator fusion without subjective weights: the region selected by the first frontier exhibits the best overall performance across all four indicators, and no single region outperforms it in all indicators. Regions with large crowding distances within the first frontier are located at the edge or sparse areas of the indicator space, corresponding to the vicinity of aero-engine performance boundaries (such as surge boundary lines, thermal load limit lines, and stall boundary lines). These regions are precisely the most risky and critical areas requiring the most thorough verification in digital equipment testing. The key areas identified by this method can scientifically guide the allocation of test bench resources, computing resources, and test cycles, enabling the completion of core tasks such as compressor surge boundary assessment, turbine blade thermal load limit testing, and overall machine performance margin assessment at the lowest test cost.

[0062] In one embodiment, Latin hypercube sampling is used to generate sample points, constructing a sample set containing sample points and corresponding response values, including: exist d In the dimensional experimental sample space, Latin hypercube sampling is used to generate sample points, and each sample point contains d The coordinate values ​​of each dimension are used to calculate the response value corresponding to each sample point through the experimental objective function or black box response function, forming a sample set, where each row of the sample set represents a sample point and its response value.

[0063] Specifically, Latin hypercube sampling is a hierarchical sampling technique that ensures uniform coverage of samples in high-dimensional space. Sample points generated through Latin hypercube sampling are divided into equally probable intervals in each dimension, and a sample point is randomly selected within each interval, thus ensuring the representativeness of the samples across the entire experimental sample space. After the sample set is constructed, the response value corresponding to each sample point is calculated using the experimental objective function or black-box response function, providing a data foundation for subsequent calculations of four-dimensional characterization indicators.

[0064] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but may be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0065] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0066] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A key region identification method based on four-dimensional representation index construction, characterized in that, The method includes: In the test sample space of digital equipment testing, sample points are generated by Latin hypercube sampling. Each sample point is calculated to obtain a corresponding response value through the engine performance model driven by digital space. The response value is at least one of the aero-engine performance indicators. A sample set containing sample points and corresponding response values ​​is constructed. Each dimension of the test sample space is composed of at least one of the aero-engine physical parameters. Based on the sample set, the information surface entropy, maximum gradient consistency, local-global similarity divergence, and local-global extreme point ratio are calculated to form a four-dimensional characterization index vector of the key region. The four-dimensional characterization index vectors are fused based on Pareto dominance. Candidate key regions are divided into different frontier levels through non-dominated sorting. The crowding distance within each frontier region is calculated. Key regions in the test sample space are obtained by combining the frontier level and the crowding distance. The key regions correspond to the mapping subspace of surge boundary region, thermal load critical region, performance inflection point region or stability transition region in the test sample space.

2. The method according to claim 1, characterized in that, The calculation process of the information surface entropy includes: A walking sequence is generated based on the space filling design, and the optimal point in the sample set is determined based on the walking sequence. Construct the adaptation vector and reference vector based on the optimal point; The entropy of the information surface is calculated based on the aforementioned adaptation vector and reference vector: in, The number of sample points in the walk sequence. To adapt the first vector i The component and the first j The measurement between the components For the reference vector, the first i The component and the first j The measurement between the components.

3. The method according to claim 1, characterized in that, The calculation process for the maximum gradient consistency includes: Based on the walking sequence generated by Latin hypercube sampling, a normalized scale is defined to eliminate the influence of dimensions, and a minimum value is set to avoid the division by zero problem; The mean of the dot product of the gradient vectors of the sample points in the sample set is calculated based on the normalized scale, and the maximum gradient consistency is obtained as follows: in, For the first i The difference in response values ​​between the groups of samples. For sample points, For normalization, For the maximum response value, For the minimum response value, For the set minimum value, Step size, This represents the number of sample points in the walk sequence.

4. The method according to claim 1, characterized in that, The calculation process of the local-to-global similarity divergence includes: The search space is divided into equally probable intervals by Latin hypercube sampling, and sample points are generated in the global space and local regions respectively. Gaussian kernel density estimation is used to estimate the probability density functions of the local distribution and the global reference distribution, respectively; The local-to-global similarity divergence is calculated based on KL divergence as follows: in, For local distribution at grid points The probability density estimate at that location, This is an estimate of the probability density of the global reference distribution at the same grid point. For the first k Volume element of each grid point The total number of grid points. These are discrete grid points.

5. The method according to claim 4, characterized in that, The method of estimating the probability density functions of the local distribution and the global reference distribution using Gaussian kernel density estimation includes: The probability density estimate of the local distribution is calculated as follows: The probability density estimate of the global reference distribution is calculated as follows: in, and For different bandwidth parameters, For Gaussian kernel function, For spatial dimensions, For local area sample points, For global spatial sample points, m This represents the total number of sample points in the global space. j The index number of the sample point.

6. The method according to claim 1, characterized in that, The calculation process for the local-to-global extreme point ratio includes: The BFGS algorithm is used to search for local extrema within a local region, and the number of local extrema is counted. Count the number of local extrema within the global scope, and calculate the ratio of local to global extrema: in, Indicates a local area The number of local extrema within, Indicates global scope The number of local extrema within, That is, the local region is a subset of the global region.

7. The method according to claim 1, characterized in that, The fusion of the four-dimensional representation index vector based on Pareto dominance includes: For any two candidate key regions, compare their four-dimensional representation index vectors. If all indices of one region are not inferior to those of the other region and at least one index is superior to that of the other region, then the former is determined to dominate the latter.

8. The method according to claim 7, characterized in that, Candidate key regions are divided into different frontier levels using non-dominated ranking, including: Regions not dominated by any other region are designated as the first frontier; regions dominated only by the first frontier region are designated as the second frontier, and so on, until all candidate key regions are assigned to the corresponding frontier level.

9. The method according to claim 8, characterized in that, Calculate the crowding distance within each frontal region, and combine the frontal level with the crowding distance to screen out key regions in the experimental sample space, including: For each region within the same frontier, its crowding distance in the index space is calculated as follows: in, For the region Crowded distance, For the first j Under these indicators, the region The value of , For the first j The maximum value of each indicator. For the first j The minimum value of each indicator. For the j-th indicator, the region The value of , To exclude the frontier The two boundary areas at the beginning and end; The key regions in the experimental sample space obtained by combining the frontier hierarchy and crowding distance screening include: Prioritize regions within the first frontier as candidate key regions; Within the first frontier, areas are sorted from largest to smallest based on their crowding distance, and the areas with the largest crowding distance are selected as the final key areas.

10. The method according to claim 1, characterized in that, Latin hypercube sampling is used to generate sample points, and a sample set containing sample points and their corresponding response values ​​is constructed, including: exist d In the dimensional experimental sample space, Latin hypercube sampling is used to generate sample points, and each sample point contains d The coordinate values ​​of each dimension are used to calculate the response value corresponding to each sample point through the experimental objective function or black box response function, forming a sample set, where each row of the sample set represents a sample point and its response value.