Additive manufacturing workpiece fatigue performance prediction method based on small sample machine learning

By employing a small-sample machine learning approach and utilizing multi-source uncertainty data amplification and feature selection strategies, a fatigue performance prediction model for additively manufactured titanium alloy parts was established. This model solves the problems of time-consuming and labor-intensive traditional methods, achieving low-cost, rapid, and accurate fatigue performance prediction, and is applicable to the aerospace and defense industries.

CN121503201APending Publication Date: 2026-02-10UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511474718.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In predicting the fatigue performance of additively manufactured titanium alloy parts, traditional methods are time-consuming, labor-intensive, and difficult to obtain accurate results. In particular, under conditions of small sample data, machine learning models have low accuracy and poor generalization ability, making it difficult to achieve accurate fatigue performance prediction.

Method used

A high-quality fatigue performance augmentation dataset was constructed using a few-sample machine learning approach and a multi-source uncertainty data augmentation strategy. By combining Pearson correlation screening, feature importance ranking, and recursive elimination, a key feature-fatigue performance prediction model was established and optimized. The prediction was then performed using a random forest model.

Benefits of technology

It enables low-cost, rapid, and accurate prediction of fatigue properties of additively manufactured titanium alloy parts under small sample data conditions, improving the robustness and generalization ability of the model, and is suitable for applications in aerospace, defense, and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121503201A_ABST
    Figure CN121503201A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of metal material preparation, and particularly relates to an additive manufacturing part fatigue performance prediction method based on small sample machine learning, and the method comprises the steps: building an original data set and a verification set, amplifying the original data set to establish an amplified data set; determining key features influencing the fatigue performance of the workpiece by adopting a three-step feature screening strategy; establishing a prediction model by taking the key features as input and the fatigue performance as output, and optimizing to obtain an optimized prediction model; and verifying and iteratively optimizing the prediction performance of the optimized prediction model by adopting a verification set to obtain a fatigue performance prediction value of the workpiece. According to the prediction method, the problems of low accuracy and poor generalization ability of a machine learning model under the background of small sample data and complex relations can be solved, and precise prediction of the fatigue performance of the additive manufacturing titanium alloy based on the nearly spherical powder is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of metal material preparation, and particularly relates to a fatigue performance prediction method for additive manufacturing parts based on small sample machine learning. BACKGROUND

[0002] Additive manufacturing titanium alloy has broad application prospects in the fields of aerospace, national defense and military industry, biological medicine, etc. However, the raw material used in additive manufacturing titanium alloy is basically spherical titanium powder prepared by atomization method. The powder inevitably contains hollow powder, which will cause defects such as shrinkage cavity, shrinkage porosity and even cracks in the internal part of the part during the layer-by-layer preparation of additive manufacturing, and aggravate the risk of part failure. At the same time, the spherical titanium powder is expensive, which greatly increases the manufacturing cost of the formed part. In recent years, irregular hydrogenation-dehydrogenation (HDH) titanium powder with internal compactness and low price has been modified to near-spherical shape by high-temperature ball milling and other methods, realizing the low-cost development of titanium powder for additive manufacturing. In addition, aerospace parts usually work under high-frequency vibration cyclic stress conditions and need to meet the service life requirement of up to several decades. Fatigue performance is a key indicator for the reliability and safety design of key parts of aircraft. Therefore, accurately predicting the fatigue performance of near-spherical powder additive manufacturing titanium parts is of great importance to their wide application in the fields of aerospace, national defense and military industry, etc.

[0003] In the additive manufacturing process, there is a complex interaction between the powder physical parameters (such as average particle size, sphericity, etc.) and the process parameters (such as laser power, scanning speed, etc.), which together affect the defect size, shape and distribution, microstructure and residual stress of the part, and then determine its fatigue performance. However, due to the numerous influencing factors and complex interaction, the traditional experimental trial-and-error method and finite element method not only consume time and effort, but also have high cost, and it is difficult to obtain accurate results. In contrast, the data-driven machine learning method can mine complex nonlinear relationships from existing data without relying on a large number of experiments and complex physical models, and can make extrapolation prediction, so it has higher accuracy, flexibility and applicability, and is widely used in the field of additive manufacturing.

[0004] However, data is the basis of machine learning, and its quality and quantity directly affect the performance of the model. It is difficult to build a robust machine learning model under the condition of small sample data. The fatigue performance data of near-spherical powder additive manufacturing titanium alloy has the characteristics of limited historical data and much noisy data, and the process of obtaining new data is time-consuming and costly. Therefore, it is urgent to develop a small sample machine learning method for fatigue performance prediction of additive manufacturing parts. SUMMARY

[0005] The purpose of this invention is to provide a method for predicting the fatigue performance of additively manufactured parts based on small-sample machine learning. This prediction method can solve at least one of the problems of low accuracy and poor generalization ability of machine learning models under conditions of small sample data and complex relationships, achieving accurate prediction of the fatigue performance of additively manufactured titanium alloys based on near-spherical powder. The prediction method in this invention is based on a small-sample data amplification strategy for multi-source uncertainty. This strategy is used to construct a high-quality fatigue performance amplification dataset that is more closely related to the actual manufacturing process, realizing the development of a robust fatigue performance prediction model. This provides technical support for the widespread application of additively manufactured titanium alloys based on near-spherical powder, and also provides a reference for constructing models with strong generalization ability under conditions of small sample data.

[0006] This invention provides a method for predicting the fatigue performance of additively manufactured parts based on few-sample machine learning, which includes the following steps: S1, establish an original dataset and a validation set; the original dataset and the validation set each independently include multiple sets of sample data, and each set of sample data independently includes powder physical property parameters, process parameters, part fatigue performance test conditions, room temperature tensile property data as input features, and fatigue performance data as output features; the data in the original dataset is different from the data in the validation set; S2, augment the original dataset to create an augmented dataset; S3, a three-step feature screening strategy is used to determine the key features that affect the fatigue performance of the part; S4. Using key features as input and fatigue performance as output, establish a key feature-fatigue performance prediction model and optimize it to obtain an optimized key feature-fatigue performance prediction model. S5, the validation set is used to verify and iteratively optimize the prediction performance of the optimized key feature-fatigue performance prediction model to obtain the predicted value of the fatigue performance of the part.

[0007] In some embodiments, in step S1, the powder physical properties include at least one of sphericity, flowability, loose density, tapped density, and average particle size.

[0008] In some embodiments, the process parameters include at least one of laser power, scanning speed, scanning spacing, layer thickness, laser line energy density, and laser volume energy density.

[0009] In some embodiments, the fatigue performance test conditions of the part include at least one of stress amplitude, stress ratio, and loading frequency.

[0010] In some embodiments, the room temperature tensile properties include at least one of tensile strength, yield strength, and elongation.

[0011] In some embodiments, the fatigue performance includes at least one of fatigue limit, fatigue life, and crack propagation rate.

[0012] In some embodiments, in step S2, the construction of the amplified dataset is based on the median absolute deviation method of the mean cross-linking of multi-source uncertainty, expressed as AVG±CF×MAD, where AVG is the characteristic mean reflecting the trend of the data center in the sample, CF is the control coefficient for regulating the magnitude of data amplification, and MAD is the robust statistic median absolute deviation that measures the degree of data dispersion.

[0013] In some embodiments, the construction of the augmented dataset includes the following steps: S2-1, CF×MAD is added to or subtracted from each feature in the original dataset to augment the data; S2-2, Treat the input features and output features of each group of sample data in the original dataset as a whole and transform them to amplify each group of sample data; S2-3, finally obtaining the augmented dataset containing the original data and the augmented data.

[0014] In some embodiments, step S3, the three-step feature screening strategy includes sequentially performing Pearson correlation screening, feature importance ranking, and recursive elimination.

[0015] In some embodiments, step S3, the three-step feature screening strategy includes the following steps: S3-1, Pearson correlation screening: Calculate the Pearson correlation coefficient r between any two features used as input in the amplified dataset. If r is greater than 0.95, then the two features are strongly correlated. S3-2, Feature Importance Ranking: Based on the amplified dataset, using powder physical property parameters, process parameters, part fatigue performance test conditions, and room temperature tensile properties as input features, and fatigue performance as output features, a random forest prediction model is established to obtain the importance ranking of each input feature on fatigue performance. The feature data with lower importance among any two strongly correlated input features in S3-1 is removed, and finally n input features are retained. S3-3, Recursive Elimination: Using the n input features retained in S3-2 as the initial input feature set, remove one input feature at a time, and use the remaining input features as input and fatigue performance as output to build corresponding random forest models. Calculate the prediction error of each model until the model's minimum error no longer decreases, at which point the recursive process stops. The remaining input features after recursive elimination are the key features.

[0016] In some embodiments, the model error evaluation index used in step S3 is the coefficient of determination. R2 ,in, R 2 The value is between 0 and 1, and the closer it is to 1, the higher the accuracy of the model.

[0017] In some embodiments, step S4, the establishment of the optimized key feature-fatigue performance prediction model includes the following steps: S4-1, Establish a dataset containing key features and fatigue performance data; S4-2, the dataset is randomly divided into a training set and a test set, wherein the amount of data in the training set accounts for 70% to 80% of the total amount of data in the dataset, the amount of data in the test set accounts for 20% to 30% of the total amount of data in the dataset, and the sum of the proportions of the training set and the test set to the total amount of data is always 1. S4-3, A key feature-fatigue performance random forest prediction model is established using the training set, and the model hyperparameters are optimized using grid search and cross-validation algorithms to obtain an optimized key feature-fatigue performance prediction model. S4-4, Evaluate the predictive performance of the optimized key feature-fatigue performance prediction model using the test set.

[0018] In some embodiments, the evaluation index used in step S4-4 includes the coefficient of determination. R 2 and mean absolute error MAE ,in, R 2 The value is between 0 and 1, and the closer it is to 1, the higher the model accuracy. MAE The smaller the value, the higher the model accuracy.

[0019] In some embodiments, step S5, model prediction performance verification and iterative optimization, includes the following steps: S5-1, Input the input features of each group in the validation set into the optimized key feature-fatigue performance prediction model to obtain the fatigue performance prediction data of each group; S5-2, calculate the relative error between each set of fatigue performance data in the validation set and the corresponding set of fatigue performance prediction data. If the relative error is >10%, add the data in the validation set as training data to the training set, and repeat steps S4 to S5 until the relative error is ≤10% to obtain the predicted value of the part's fatigue performance.

[0020] The prediction method in this invention has the following advantages: 1) The fatigue performance prediction method for additively manufactured parts based on small sample machine learning proposed in this invention, compared with the traditional experimental trial and error method and the finite element analysis method based on physical models, can make full use of existing data to explore the implicit relationship between the fatigue performance of additively manufactured titanium alloys based on near-spherical powder and the influencing factors such as powder characteristics and process parameters, thereby achieving low-cost, rapid and accurate prediction of fatigue performance.

[0021] 2) This invention proposes a small sample data augmentation strategy based on multi-source uncertainty. This strategy integrates the average value (AVG) that reflects the trend of data centers with the robust statistic of median absolute deviation (MAD) that measures the degree of data dispersion. It systematically quantifies and integrates multi-source uncertainties such as powder characteristics, equipment condition fluctuations, and test errors, and constructs a high-quality dataset that is more in line with reality, providing a data foundation for machine learning models.

[0022] 3) This invention introduces a control coefficient (CF) that can adjust the amplification magnitude. By dynamically adjusting the CF value according to the characteristics of different datasets, it can effectively avoid data distribution distortion caused by overly aggressive or conservative amplification, greatly improving the method's versatility. Furthermore, this method does not require a complex model training process, resulting in higher data generation efficiency and lower computational cost.

[0023] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0024] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a fatigue performance prediction method for additively manufactured parts based on few-sample machine learning in an embodiment of the present invention. Detailed Implementation

[0026] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the invention and to fully convey the scope of the invention to those skilled in the art.

[0027] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0028] In the description of the embodiments of this invention, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this invention, "multiple" means two or more, unless otherwise explicitly defined.

[0029] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0030] In the description of the embodiments of this invention, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0031] In the description of the embodiments of the present invention, the term "multiple" refers to two or more (including two), similarly, "multiple groups" refers to two or more (including two groups), and "multiple pieces" refers to two or more (including two pieces).

[0032] In the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in the embodiments of the present invention according to the specific circumstances.

[0033] See Figure 1 As shown, the fatigue performance prediction method for additively manufactured parts based on few-sample machine learning provided by this invention is carried out according to the following steps.

[0034] S1, Create the original dataset and validation set.

[0035] In this embodiment of the invention, the original dataset and the validation set each independently include multiple sets of sample data, and each set of sample data independently includes powder physical property parameters, process parameters, part fatigue performance test conditions, room temperature tensile performance data as input features, and fatigue performance data as output features; moreover, the data in the original dataset is different from the data in the validation set. It can be understood that the original dataset and the validation set are two different and independent datasets.

[0036] In some embodiments, a raw dataset is established based on historical data of near-spherical powder additive manufacturing of titanium alloys, including data collected from experiments, published literature and patents, which includes powder physical property parameters, process parameters, part fatigue performance test conditions, room temperature tensile property data and fatigue performance data.

[0037] In embodiments of the present invention, the original dataset includes multiple sets of sample data, and each set of sample data independently includes powder physical property parameters, process parameters, part fatigue performance test conditions, room temperature tensile performance data, and fatigue performance data; wherein, powder physical property parameters, process parameters, part fatigue performance test conditions, and room temperature tensile performance are used as input features, and fatigue performance is used as output features.

[0038] As an embodiment of the present invention, the powder physical property parameters in the original dataset include at least one of sphericity, flowability, loose density, tap density, and average particle size.

[0039] As an embodiment of the present invention, the process parameters in the original data set include at least one of laser power, scanning speed, scanning spacing, layer thickness, laser line energy density, and laser volume energy density.

[0040] As an embodiment of the present invention, the fatigue performance test conditions of the original data include at least one of stress amplitude, stress ratio, and loading frequency.

[0041] As an embodiment of the present invention, the room temperature tensile properties in the original dataset include at least one of tensile strength, yield strength, and elongation.

[0042] As an embodiment of the present invention, the fatigue performance of the original dataset includes at least one of fatigue limit, fatigue life, and crack propagation rate.

[0043] In some embodiments, 5 to 10 sets of fatigue performance data of near-spherical powder additive manufacturing of titanium alloys are collected to establish a validation set for model validation, and it is ensured that the data in the validation set are different from the data in the original dataset.

[0044] Specifically, the validation set includes multiple sets of sample data, and each set of sample data independently includes powder physical property parameters, process parameters, part fatigue performance test conditions, room temperature tensile performance data, and fatigue performance data. Among them, powder physical property parameters, process parameters, part fatigue performance test conditions, and room temperature tensile performance are used as input features, and fatigue performance is used as output features. The limitations of the input and output features are the same as those in the original dataset.

[0045] As an embodiment of the present invention, the verified powder physical properties include at least one of sphericity, flowability, loose density, tapped density, and average particle size.

[0046] As an embodiment of the present invention, the verification of the process parameters includes at least one of laser power, scanning speed, scanning spacing, layer thickness, laser line energy density, and laser volume energy density.

[0047] As an embodiment of the present invention, the test conditions for verifying the fatigue performance of the concentrated component include at least one of stress amplitude, stress ratio, and loading frequency.

[0048] As an embodiment of the present invention, the verification of concentrated room temperature tensile properties includes at least one of tensile strength, yield strength, and elongation.

[0049] As an embodiment of the present invention, the verification of concentrated fatigue performance includes at least one of fatigue limit, fatigue life, and crack propagation rate.

[0050] S2, Create an expanded dataset.

[0051] In this embodiment of the invention, the original dataset is augmented to create an augmented dataset.

[0052] In this embodiment of the invention, the original dataset is amplified based on the average cross-linked median absolute deviation method (AVG±CF×MAD) for multi-source uncertainty to establish an amplified dataset; wherein, AVG is the characteristic average value reflecting the data center trend of the sample, CF is the control coefficient for regulating the data amplification magnitude, and MAD is the robust statistic median absolute deviation that measures the degree of data dispersion.

[0053] In this embodiment of the invention, the construction of the augmented dataset specifically includes the following steps: S2-1 augments the data by adding or subtracting CF×MAD from each feature in the original dataset. Specifically, for each feature, CF×MAD is added or subtracted independently, while keeping the other features unchanged. Assume each set of samples in the original dataset contains m features (the sum of input and output features), and each feature corresponds to two transformations (adding CF×MAD and subtracting CF×MAD). Therefore, each set of samples generates 2^m new sets of data, or 2^m augmented sets of data. If the original dataset contains k sets of samples, then 2^m × k new sets of data are generated, or 2^m × k augmented sets of data.

[0054] S2-2 transforms each sample data set in the original dataset by treating its input and output features as a whole. Specifically, for each sample data set, CF×MAD is added or subtracted independently from each input feature or each output feature. Then, the input features, input features minus CF×MAD, and input features plus CF×MAD are randomly combined with the output features, output features minus CF×MAD, and output features plus CF×MAD to obtain new sample data sets, each consisting of both input and output features. Through this process, each sample data set generates 8 new data sets, or 8 augmented data sets. The k original sample data sets generate a total of 8k new data sets, or 8k augmented data sets.

[0055] S2-3, finally obtaining the augmented dataset containing the original data and the amplified data. Specifically, the final result is an augmented dataset containing k sets of original data and (2 m×k+8 k) sets of new data.

[0056] It is worth mentioning that the method of average cross-median absolute deviation (AVG±CF×MAD) based on multi-source uncertainty amplifies the original dataset by using a cross-combination approach rather than simple addition or subtraction. This approach can better maintain the diversity and representativeness of the data and generate a more realistic and high-quality dataset. At the same time, the CF value can be dynamically adjusted according to the characteristics of the dataset, thereby effectively avoiding overly aggressive or conservative amplification and improving the method's wide applicability.

[0057] S3, Feature Filtering.

[0058] In this embodiment of the invention, a three-step feature screening strategy is used to determine the key features that affect the fatigue performance of the part.

[0059] In this embodiment of the invention, the three-step feature screening strategy includes sequentially performing Pearson correlation screening, feature importance ranking, and recursive elimination. That is, the three-step feature screening strategy of “Pearson correlation screening - feature importance ranking - recursive elimination” is used to determine the key features that affect the fatigue performance of the part.

[0060] As an embodiment of the present invention, the three-step feature screening strategy specifically includes the following steps: S3-1, Pearson correlation screening: Based on the augmented dataset, calculate the Pearson correlation coefficient (r) between any two features used as input, as shown in formula (1).

[0061] (1) In the formula, The correlation coefficient is... and Represents any two input features, Represents the covariance function. and It represents the standard deviation.

[0062] Specifically, when r ≤ 0.95, the two input features are not strongly correlated and can both be used as independent variables as model input features.

[0063] When r>0.95, there is a strong correlation between the two input features, and it is necessary to remove one of the features that has a smaller impact on fatigue performance by comparison.

[0064] S3-2, Feature Importance Ranking: Based on the amplified dataset, using powder physical property parameters, process parameters, part fatigue performance test conditions, and room temperature tensile properties as input features, and fatigue performance as output features, a random forest prediction model is established to obtain the importance ranking of each input feature on fatigue performance. The feature data with lower importance among any two strongly correlated input features in step S3-1 is removed, and finally n input features are retained.

[0065] In the embodiments of the present invention, when modeling the random forest prediction model, the augmented dataset is divided into a training set and a test set. The training set and the test set each independently contain the features as input and the fatigue performance as output. The training set is used for model training, and its data volume accounts for 70% to 80% of the total data volume of the augmented dataset. The test set is used to evaluate the model's prediction performance, and its data volume accounts for 20% to 30% of the total data volume of the augmented dataset. The sum of the proportions of the training set and the test set to the total data volume of the augmented dataset is always 1.

[0066] S3-3, Recursive Elimination: Using the n input features retained in S3-2 as the initial input feature set, first remove one of the input features, and use the remaining n-1 input features as input and fatigue performance as output to build the corresponding random forest model; then remove one of the n-1 input features, and use the remaining n-2 input features as input and fatigue performance as output to build the corresponding random forest model; then repeat the above process with the remaining n-2 input features, and so on, calculating the prediction error of each model until the minimum error of the model no longer decreases, at which point the recursive process stops, and the remaining input features after recursive elimination are the key features.

[0067] The model error evaluation index used in this embodiment of the invention is the coefficient of determination. R 2 As shown in formula (2), its value is in the range of 0 to 1, and the closer it is to 1, the higher the accuracy of the model.

[0068] (2) In the formula, and These represent the predicted and actual fatigue performance values, respectively. This represents the average value of the actual fatigue performance. This is to increase the number of samples in the dataset. It can be understood that each sample corresponds to a set of input features and output fatigue performance data.

[0069] S4. Establish a key feature-fatigue performance prediction model.

[0070] In this embodiment of the invention, a key feature-fatigue performance prediction model is established using key features as input and fatigue performance as output, and then optimized to obtain an optimized key feature-fatigue performance prediction model.

[0071] As an embodiment of the present invention, the construction of the key feature-fatigue performance prediction model includes the following steps: S4-1, Based on the key features and fatigue performance generated in step S3, establish a dataset containing key features and fatigue performance data.

[0072] S4-2, randomly divide the dataset from step S4-1 into a training set and a test set. The training set accounts for 70% to 80% of the total dataset, and the test set accounts for 20% to 30% of the total dataset. The sum of the proportions of the training set and the test set to the total dataset is always 1.

[0073] S4-3. Using the training set described in step S4-2, establish a key feature-fatigue performance random forest prediction model, and use grid search combined with K-Fold Cross Validation algorithm to optimize the model hyperparameters to obtain an optimized key feature-fatigue performance prediction model.

[0074] S4-4, Use the test set described in step S4-2 to evaluate the predictive performance of the optimized key feature-fatigue performance prediction model described in step S4-3.

[0075] In this embodiment of the invention, the evaluation index used to evaluate the predictive performance of the optimized key feature—fatigue performance prediction model—is the coefficient of determination. R 2 and mean absolute error MAE The calculation formulas are shown in (2) to (3): (3) In the formula, and These represent the predicted and actual fatigue performance values, respectively. This represents the average value of the actual fatigue performance. This represents the number of samples. It can be understood that each sample corresponds to a set of key features as input and fatigue performance data as output.

[0076] Specifically, R 2 The value is in the range of 0 to 1, and the closer it is to 1, the higher the accuracy of the model; MAE The smaller the value, the higher the model accuracy.

[0077] S5, Model Performance Verification and Iterative Optimization.

[0078] In this embodiment of the invention, a validation set is used to verify and iteratively optimize the prediction performance of the optimized key feature-fatigue performance prediction model to obtain the predicted value of the fatigue performance of the part.

[0079] In some embodiments, 5 to 10 sets of fatigue performance data of near-spherical powder additive manufacturing of titanium alloys are collected to establish a validation set for model validation, and it is ensured that the data in the validation set are different from the data in the original dataset.

[0080] In this embodiment of the invention, the model prediction performance verification and iterative optimization include the following steps: S5-1, input the input features of each group in the validation set into the optimized key feature-fatigue performance prediction model to obtain the fatigue performance prediction data for each group. It can be understood that in the validation set, each group of input features corresponds to a set of fatigue performance data (hereinafter referred to as actual values); simultaneously, after each group of input features is input into the optimized key feature-fatigue performance prediction model, it outputs a set of fatigue performance prediction data, i.e., fatigue performance prediction values ​​(hereinafter referred to as predicted values). Thus, each set of actual values ​​corresponds to a set of predicted values.

[0081] S5-2 Calculate the relative error between each set of fatigue performance data (actual value) and the corresponding set of fatigue performance prediction data (predicted value) in the verification set. When the relative error between the actual value and the predicted value is ≤10%, the construction of the best key feature-fatigue performance prediction model is completed, and the predicted value of the fatigue performance of the part is obtained simultaneously.

[0082] In some embodiments, if the relative error between the actual value and the predicted value is >10%, the data in the validation set is added to the training set as training data, and steps S4 to S5 are repeated until the relative error between the predicted value and the actual value is ≤10%, thus completing the construction of the best key feature-fatigue performance prediction model and simultaneously obtaining the predicted value of the part's fatigue performance.

[0083] In some embodiments, the predicted values ​​from the group with the smallest relative error are used as the predicted fatigue performance values.

[0084] Unless otherwise defined, the technical terms used in the following embodiments have the same meaning as commonly understood by those skilled in the art. Unless otherwise specified, the experimental reagents used in the following embodiments are all conventional biochemical reagents; the raw materials, instruments, and equipment used in the following embodiments can all be obtained commercially or through existing methods; unless otherwise specified, the amounts of experimental reagents used are the amounts used in conventional experimental operations; unless otherwise specified, the experimental methods are conventional methods. It should be further noted that the following description is merely exemplary and not a specific limitation of the present invention.

[0085] Example 1 Taking the fatigue life prediction of Ti-6Al-4V alloy based on near-spherical powder additive manufacturing as an example, this paper details the implementation process of the fatigue performance prediction method for additively manufactured parts based on data augmentation strategies. The specific steps of the fatigue performance prediction method for additively manufactured parts based on small-sample machine learning are as follows: 1) Establish the original dataset: Collect data on powder sphericity, laser power, scanning speed, scanning spacing, laser volume energy density, stress amplitude, stress ratio, part yield strength, elongation and fatigue life from experiments, public literature and patents, and establish an original dataset containing 30 sets of data; among them, part fatigue life is used as the output feature, and the other features are used as the input features.

[0086] 2) Establishing the augmented dataset: The original dataset was augmented using the average cross-linked median absolute deviation method (AVG±CF×MAD) based on multi-source uncertainty. The original dataset contained 30 sets of data, each set including 9 input features and 1 output feature, resulting in 840 new sets of data. Finally, an augmented dataset containing 870 data points was established.

[0087] 3) Feature selection: Key features affecting fatigue life were identified through a progressive process of Pearson correlation screening, feature importance ranking based on random forest, and recursive elimination. Ultimately, powder sphericity, laser energy density, stress amplitude, stress ratio, part yield strength, and elongation were determined as key features for the fatigue life prediction model.

[0088] 4) Establishing a Key Feature-Fatigue Life Prediction Model: First, based on the feature selection results, a dataset containing key features and fatigue life is established for building the prediction model. The dataset is randomly divided in an 8:2 ratio, with 80% of the total data used as the training set for training the fatigue life prediction model, and the remaining 20% ​​used as the test set for evaluating model prediction performance. During model training, grid search combined with cross-validation is used to optimize the model's hyperparameters, obtaining an optimized key feature-fatigue life prediction model. The coefficient of determination R0 of the optimized key feature-fatigue life prediction model is... 2 The mean absolute error (MAE) is 0.917 and the mean absolute error (MAE) is 0.215.

[0089] 5) Model Performance Validation and Iterative Optimization: First, five sets of new data, different from the original dataset, were collected from previous experiments, published literature, and patents to establish a validation set. Each set of data in the validation set independently includes the input features for modeling the key features described in step 3) (i.e., step 4), specifically powder sphericity, laser energy density, stress amplitude, stress ratio, part yield strength, elongation, and fatigue performance data. Then, the input features of each set of data in the validation set were input into the optimized key feature-fatigue life prediction model to obtain the corresponding fatigue life prediction values.

[0090] The relative error between the predicted values ​​and the actual values ​​in the validation set is calculated, as shown in Table 1.

[0091] Table 1

[0092] The results show that the relative errors of the five sets of data in Example 1 are all ≤10%, indicating that the constructed fatigue performance prediction model meets the usage requirements. Furthermore, the predicted value from the set with the smallest relative error is used as the fatigue life prediction value in Example 1.

[0093] Example 2 Taking the fatigue life prediction of pure titanium additive manufacturing based on near-spherical powder as an example, this paper details the implementation process of the fatigue performance prediction method for additive manufacturing parts incorporating data augmentation strategies. The specific steps of the fatigue performance prediction method for additive manufacturing parts based on small-sample machine learning are as follows: 1) Establish the original dataset: Collect data on powder sphericity, laser power, scanning speed, scanning spacing, stress amplitude, stress ratio, tensile strength, yield strength, elongation and fatigue life of the parts from experiments, public literature and patents, and establish an original dataset containing 25 sets of data; among them, the fatigue life of the parts is used as the output feature, and the other features are used as the input features.

[0094] 2) Establishing the augmented dataset: The original dataset was augmented using the average cross-linked median absolute deviation method (AVG±CF×MAD) based on multi-source uncertainty. The original dataset contained 25 sets of data, each set including 9 input features and 1 output feature, thus obtaining 700 new sets of data. Finally, an augmented dataset containing 725 data points was established.

[0095] 3) Feature selection: Key features affecting fatigue life were identified through a series of methods, including Pearson correlation screening, random forest-based feature importance ranking, and recursive elimination. Ultimately, powder sphericity, laser power, scanning speed, stress amplitude, stress ratio, tensile strength of the workpiece, and yield strength were determined as the key features for the fatigue life prediction model.

[0096] 4) Establishing a Key Feature-Fatigue Life Prediction Model: First, based on the feature selection results, a dataset containing key features and fatigue life is established for building the prediction model. The dataset is randomly divided in a 7:3 ratio, with 70% of the total data used as the training set for training the fatigue life prediction model, and the remaining 30% used as the test set for evaluating model prediction performance. During model training, grid search combined with cross-validation is used to optimize the model's hyperparameters, obtaining an optimized key feature-fatigue life prediction model. The coefficient of determination R of the optimized fatigue life prediction model is... 2 The mean absolute error (MAE) is 0.912 and the mean absolute error (MAE) is 0.225.

[0097] 5) Model Performance Validation and Iterative Optimization: First, five sets of new data, different from the original dataset, were collected from previous experiments, published literature, and patents to establish a validation set. Each set of data in the validation set independently includes the key features described in step 3) (i.e., the input features for modeling in step 4, specifically powder sphericity, laser power, scanning speed, stress amplitude, stress ratio, tensile strength of the part, and yield strength) and fatigue life data. Then, the input features of each set of data in the validation set were input into the optimized key feature-fatigue life prediction model to obtain the corresponding fatigue life prediction values.

[0098] The relative error between the predicted values ​​and the actual values ​​in the validation set is calculated, as shown in Table 2.

[0099] Table 2

[0100] The results show that the relative errors of all five sets of data in Example 2 are ≤10%, indicating that the constructed fatigue performance prediction model meets the usage requirements. Furthermore, the predicted value from the set with the smallest relative error is used as the fatigue life prediction value in Example 2.

[0101] Example 3 Taking the fatigue life prediction of pure titanium additive manufacturing based on near-spherical powder as an example, this paper details the implementation process of the fatigue performance prediction method for additive manufacturing parts incorporating data augmentation strategies. The specific steps of the fatigue performance prediction method for additive manufacturing parts based on small-sample machine learning are as follows: 1) Establish the original dataset: Collect data on powder flowability, laser power, scanning speed, scanning spacing, laser volume energy density, stress amplitude, stress ratio, part yield strength and fatigue life from experiments, public literature and patents, and establish an original dataset containing 30 sets of data; among them, part fatigue life is used as the output feature, and the other features are used as the input features.

[0102] 2) Establishing the augmented dataset: The original dataset was augmented using the average cross-linked median absolute deviation method (AVG±CF×MAD) based on multi-source uncertainty. The original dataset contained 30 sets of data, each set including 8 input features and 1 output feature, thus obtaining 780 new sets of data. Finally, an augmented dataset containing 810 data points was established.

[0103] 3) Feature Selection: Key features affecting fatigue life were identified through a progressive process involving Pearson correlation screening, random forest-based feature importance ranking, and recursive elimination. Ultimately, powder flowability, laser energy density, stress amplitude, stress ratio, and part yield strength were determined as the key features for the fatigue life prediction model.

[0104] 4) Establishing a Key Feature-Fatigue Life Prediction Model: First, based on the feature selection results, a dataset containing key features and fatigue life is established for building the prediction model. The dataset is randomly divided in a 7:3 ratio, with 70% of the total data used as the training set for training the fatigue life prediction model, and the remaining 30% used as the test set for evaluating the model's prediction performance. During model training, grid search combined with cross-validation is used to optimize the model's hyperparameters, obtaining an optimized key feature-fatigue life prediction model. The coefficient of determination R0 of the optimized key feature-fatigue life prediction model is... 2 The mean absolute error (MAE) is 0.904 and the mean absolute error (MAE) is 0.240.

[0105] 5) Model Performance Validation and Iterative Optimization: First, five sets of new data, different from the original dataset, were collected from previous experiments, published literature, and patents to establish a validation set. Each set of data in the validation set independently includes the input features for modeling the key features described in step 3) (i.e., step 4), specifically powder flowability, laser energy density, stress amplitude, stress ratio, and part yield strength, as well as fatigue life data. Then, the input features of each set of data in the validation set were input into the optimized key feature-fatigue life prediction model to obtain the corresponding fatigue life prediction values.

[0106] The relative error between the predicted values ​​and the actual values ​​in the validation set is calculated, as shown in Table 3.

[0107] Table 3

[0108] The results show that the relative errors of all five sets of data in Example 3 are ≤10%, indicating that the constructed fatigue performance prediction model meets the usage requirements. Furthermore, the predicted value from the set with the smallest relative error is used as the fatigue life prediction value in Example 3.

[0109] Example 4 Taking the fatigue limit prediction of Ti-6Al-4V alloy parts based on near-spherical powder additive manufacturing as an example, this paper details the implementation process of the fatigue performance prediction method for additively manufactured parts based on data augmentation strategies. The specific steps of the fatigue performance prediction method for additively manufactured parts based on small-sample machine learning are as follows: 1) Establish the original dataset: Collect data on powder flowability, powder sphericity, laser power, scanning speed, scanning spacing, laser volume energy density, stress amplitude, stress ratio, part yield strength, tensile strength, elongation, and fatigue limit from experiments, published literature, and patents to establish an original dataset containing 20 sets of data; among them, the part fatigue limit is used as the output feature, and the other features are used as input features; the maximum stress amplitude at which the material does not fail when the number of cycles reaches 107.

[0110] 2) Establishing the augmented dataset: The original dataset was augmented using the average cross-linked median absolute deviation method (AVG±CF×MAD) based on multi-source uncertainty. The original dataset contained 20 sets of data, each set including 11 input features and 1 output feature, resulting in 640 new sets of data. Finally, an augmented dataset containing 660 data points was established.

[0111] 3) Feature selection: Key features affecting fatigue life were identified through a series of methods, including Pearson correlation screening, random forest-based feature importance ranking, and recursive elimination. Ultimately, powder flowability, laser energy density, stress amplitude, stress ratio, tensile strength of the workpiece, and yield strength were determined as the key features for the fatigue limit prediction model.

[0112] 4) Establishing a Key Feature-Fatigue Limit Prediction Model: First, based on the feature selection results, a dataset containing key features and fatigue limits is established for building the prediction model. The dataset is randomly divided in an 8:2 ratio, with 80% of the total data used as the training set for training the fatigue limit prediction model, and the remaining 20% ​​used as the test set for evaluating model prediction performance. During model training, grid search combined with cross-validation is used to optimize the model's hyperparameters, obtaining an optimized key feature-fatigue limit prediction model. The coefficient of determination R0 of the optimized key feature-fatigue limit prediction model is... 2 The mean absolute error (MAE) is 0.925 and the mean absolute error (MAE) is 0.194.

[0113] 5) Model Performance Validation and Iterative Optimization: First, eight sets of new data, different from the original dataset, were collected from previous experiments, published literature, and patents to establish a validation set. Each set of data in the validation set independently includes the input features for modeling the key features described in step 3) (i.e., step 4), specifically powder flowability, laser energy density, stress amplitude, stress ratio, tensile strength of the part, yield strength, and fatigue limit data. Then, the input features of each set of data in the validation set were input into the optimized key feature-fatigue limit prediction model to obtain the corresponding fatigue limit prediction values.

[0114] The relative error between the predicted value and the actual value in the validation set is calculated, as shown in Table 4.

[0115] Table 4

[0116] The results show that the relative errors of all eight sets of data in Example 4 are ≤10%, indicating that the constructed fatigue performance prediction model meets the usage requirements. Furthermore, the predicted value from the set with the smallest relative error is used as the fatigue limit prediction value in Example 4.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting the fatigue performance of additively manufactured parts based on few-sample machine learning, characterized in that, Includes the following steps: S1, establish an original dataset and a validation set; the original dataset and the validation set each independently include multiple sets of sample data, and each set of sample data independently includes powder physical property parameters, process parameters, part fatigue performance test conditions, room temperature tensile property data as input features, and fatigue performance data as output features; the data in the original dataset is different from the data in the validation set; S2, augment the original dataset to create an augmented dataset; S3, a three-step feature screening strategy is used to determine the key features that affect the fatigue performance of the part; S4. Using key features as input and fatigue performance as output, establish a key feature-fatigue performance prediction model and optimize it to obtain an optimized key feature-fatigue performance prediction model. S5, the validation set is used to verify and iteratively optimize the prediction performance of the optimized key feature-fatigue performance prediction model to obtain the predicted value of the fatigue performance of the part.

2. The fatigue performance prediction method for additively manufactured parts based on few-sample machine learning as described in claim 1, characterized in that, In step S1, the powder physical property parameters include at least one of sphericity, flowability, loose density, tap density, and average particle size; Preferably, the process parameters include at least one of laser power, scanning speed, scanning spacing, layer thickness, laser line energy density, and laser volume energy density; Preferably, the fatigue performance test conditions for the part include at least one of stress amplitude, stress ratio, and loading frequency; Preferably, the room temperature tensile properties include at least one of tensile strength, yield strength, and elongation; Preferably, the fatigue performance includes at least one of fatigue limit, fatigue life, and crack propagation rate.

3. The fatigue performance prediction method for additively manufactured parts based on few-sample machine learning as described in claim 1, characterized in that, In step S2, the construction of the amplified dataset is based on the median absolute deviation method of the mean cross-linking of multi-source uncertainty, expressed as AVG±CF×MAD, where AVG is the characteristic mean reflecting the trend of the data center in the sample, CF is the control coefficient for regulating the data amplification amplitude, and MAD is the robust statistic median absolute deviation that measures the degree of data dispersion.

4. The fatigue performance prediction method for additively manufactured parts based on few-sample machine learning as described in claim 3, characterized in that, The construction of the augmented dataset includes the following steps: S2-1, CF×MAD is added to or subtracted from each feature in the original dataset to augment the data; S2-2, Treat the input features and output features of each group of sample data in the original dataset as a whole and transform them to amplify each group of sample data; S2-3, finally obtaining the augmented dataset containing the original data and the augmented data.

5. The fatigue performance prediction method for additively manufactured parts based on few-sample machine learning as described in claim 1, characterized in that, In step S3, the three-step feature screening strategy includes sequentially performing Pearson correlation screening, feature importance ranking, and recursive elimination.

6. The fatigue performance prediction method for additively manufactured parts based on few-sample machine learning as described in claim 5, characterized in that, In step S3, the three-step feature selection strategy includes the following steps: S3-1, Pearson correlation screening: Calculate the Pearson correlation coefficient r between any two features used as input in the amplified dataset. If r is greater than 0.95, then the two features are strongly correlated. S3-2, Feature Importance Ranking: Based on the amplified dataset, using powder physical property parameters, process parameters, part fatigue performance test conditions, and room temperature tensile properties as input features, and fatigue performance as output features, a random forest prediction model is established to obtain the importance ranking of each input feature on fatigue performance. The feature data with lower importance among any two strongly correlated input features in S3-1 is removed, and finally n input features are retained. S3-3, Recursive Elimination: Using the n input features retained in S3-2 as the initial input feature set, remove one input feature at a time, and use the remaining input features as input and fatigue performance as output to build corresponding random forest models. Calculate the prediction error of each model until the model's minimum error no longer decreases, at which point the recursive process stops. The remaining input features after recursive elimination are the key features.

7. The fatigue performance prediction method for additively manufactured parts based on few-sample machine learning as described in claim 6, characterized in that, In step S3, the model error evaluation index used is the coefficient of determination. R 2 ,in, R 2 The value is between 0 and 1, and the closer it is to 1, the higher the accuracy of the model.

8. The fatigue performance prediction method for additively manufactured parts based on few-sample machine learning as described in claim 1, characterized in that, In step S4, the establishment of the optimized key feature - fatigue performance prediction model includes the following steps: S4-1, Establish a dataset containing key features and fatigue performance data; S4-2, the dataset is randomly divided into a training set and a test set, wherein the amount of data in the training set accounts for 70% to 80% of the total amount of data in the dataset, the amount of data in the test set accounts for 20% to 30% of the total amount of data in the dataset, and the sum of the proportions of the training set and the test set to the total amount of data is always 1. S4-3, A key feature-fatigue performance random forest prediction model is established using the training set, and the model hyperparameters are optimized using grid search and cross-validation algorithms to obtain an optimized key feature-fatigue performance prediction model. S4-4, Evaluate the predictive performance of the optimized key feature-fatigue performance prediction model using the test set.

9. The fatigue performance prediction method for additively manufactured parts based on few-sample machine learning as described in claim 8, characterized in that, In step S4-4, the evaluation indicators used include the coefficient of determination. R 2 and mean absolute error MAE , in, R 2 The value is between 0 and 1, and the closer it is to 1, the higher the model accuracy. MAE The smaller the value, the higher the model accuracy.

10. The fatigue performance prediction method for additively manufactured parts based on few-sample machine learning as described in claim 1, characterized in that, In step S5, the model prediction performance verification and iterative optimization include the following steps: S5-1, Input the input features of each group in the validation set into the optimized key feature-fatigue performance prediction model to obtain the fatigue performance prediction data of each group; S5-2, calculate the relative error between each set of fatigue performance data in the validation set and the corresponding set of fatigue performance prediction data. If the relative error is >10%, add the data in the validation set as training data to the training set, and repeat steps S4 to S5 until the relative error is ≤10% to obtain the predicted value of the part's fatigue performance.