Factor analysis-based random forest laser absorption coefficient efficient prediction method and system

By combining factor analysis and random forest model, a laser absorption coefficient dataset was constructed and preprocessed and data incremented, which solved the problem of difficult laser absorption coefficient selection, achieved efficient laser processing numerical simulation, and improved research efficiency and accuracy.

CN120705834APending Publication Date: 2025-09-26HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510826951.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, the selection of the laser absorption coefficient is difficult, resulting in low efficiency of numerical simulation of laser processing. In particular, when process parameters or materials change, repeated adjustments are required to achieve the best simulation effect, and there is a lack of empirical support.

Method used

A random forest model based on factor analysis was adopted. By constructing a laser absorption coefficient dataset, preprocessing and data increment were performed, and factor analysis and random forest regression were used to predict the laser absorption coefficient. Factor analysis and Gaussian mixture model were combined to optimize feature selection and improve model training efficiency.

Benefits of technology

It achieves efficient prediction of laser absorption coefficient, improves the efficiency and accuracy of laser processing numerical simulation, reduces model adjustment time, and provides an efficient parameter estimation method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705834A_ABST
    Figure CN120705834A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of laser processing numerical simulation, and discloses a random forest model laser absorption coefficient efficient prediction method based on factor analysis, and a laser absorption coefficient data set comprises four parameters of laser power density, a forming coefficient, conductivity and thermal diffusivity. According to the method, multiple factors influencing the laser absorption coefficient in laser processing are comprehensively considered, the dimensionality of a data set is reduced, the feature correlation between data is low, and a high-quality training set is provided for subsequent training of a random forest model. Multiple Gaussian distribution linear mixing probability distributions are generated through a Gaussian mixture model to describe original data, new samples are generated through random sampling to form an incremental data set, the training effect of the random forest model is improved, and the over-fitting condition is avoided; through factor analysis, features with strong correlation are divided into one group, generation of a weak decision tree is avoided, the accuracy and convergence speed of the model are improved, and an efficient model training mode is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of laser processing numerical simulation, and in particular relates to a method and system for efficiently predicting laser absorption coefficient using a random forest model based on factor analysis. Background Art

[0002] Quantifying laser energy absorption is crucial for understanding laser-material interactions and melt pool dynamics. The laser absorption coefficient characterizes a medium's ability to absorb light, determining the proportion of laser energy absorbed by a metal material and directly influencing energy utilization and distribution. In the field of numerical simulation of laser processing, selecting the appropriate laser absorption coefficient is fundamental to accurately reproducing the laser energy heat input and is a key factor influencing the accuracy of numerical simulation results.

[0003] The laser absorption coefficient is determined by the material properties and the type of laser. During the laser processing process, the surface of the material is at an ultra-high temperature, the material undergoes a complex and drastic phase change, and the laser absorption coefficient is in a highly dynamic and nonlinear change process. Using the absorption coefficient at room temperature, the simulation results often have large deviations. However, obtaining the absorption coefficient at high temperatures through calorimetry or integrating sphere radiation measurement requires high equipment costs and complex experimental processes. The above factors make the selection of the laser absorption coefficient difficult. In numerical simulations, the laser absorption coefficient is usually regarded as a constant value, and simulation studies are carried out through empirical selection. However, when the process parameters or materials change, there is a lack of empirical support, and the absorption coefficient needs to be repeatedly adjusted to achieve the best simulation effect. Each adjustment requires re-simulation, which limits the efficiency of laser processing mechanism research.

[0004] Therefore, how to apply machine learning technology and use past experience data to achieve more efficient and meaningful prediction of laser absorption coefficient, provide an efficient parameter estimation method for laser processing numerical simulation technology, and improve the efficiency of mechanism research has become an urgent problem to be solved.

[0005] Through the above analysis, the problems and defects of the existing technology are as follows:

[0006] Obtaining the absorption coefficient at high temperatures through calorimetry or integrating sphere radiometry requires high equipment costs and complex experimental procedures. These factors make the selection of the laser absorption coefficient difficult. In numerical simulations, the laser absorption coefficient is usually considered a constant value, and simulation studies are carried out through empirical selection. However, when process parameters or materials change, there is a lack of empirical support, and the absorption coefficient needs to be repeatedly adjusted to achieve the best simulation effect. Each adjustment requires re-simulation, which limits the efficiency of laser processing mechanism research. Summary of the Invention

[0007] In response to the problems existing in the prior art, the present invention provides an efficient prediction method for laser absorption coefficient using a random forest model based on factor analysis.

[0008] The present invention is implemented as follows: a method for efficiently predicting laser absorption coefficient using a random forest model based on factor analysis includes:

[0009] Step 1: Establish a laser absorption coefficient dataset using past experiments, simulations, and relevant literature data;

[0010] The data set was subjected to factor analysis to select the laser energy density, forming factor, electrical conductivity, thermal diffusivity and laser absorption coefficient under different processes and welding materials;

[0011] Step 2: preprocess the data set;

[0012] This includes operations such as completion and standardization. The conductivity and thermal diffusivity of different materials in the dataset may be missing, and similar metal data are used to complete the missing data. The dataset is also Z-score standardized.

[0013] Step 3: Build a data increment model based on the Gaussian mixture model, generate a probability distribution of multiple Gaussian distribution linear mixtures to describe the original data, and randomly generate new samples through the distribution model to form an incremental data set;

[0014] Step 4: Construct a random forest laser absorption coefficient prediction model fused with factor analysis;

[0015] The laser absorption coefficient prediction model was trained using incremental data sets. Bootstrap was used to randomly extract k data sets with replacement. If k was 100, approximately 37% of the data was not selected and became the out-of-bag data, or test set. Fusion factor analysis was used to construct feature groups on the incremental data sets. After training, an integrated model consisting of multiple decision trees was obtained for predicting the laser absorption coefficient.

[0016] In step 5, the forming coefficient of the laser-processed product is obtained through experiments. At the same time, the laser energy density is calculated based on the laser power, moving speed and spot size of the laser used, and a set of data is formed by combining the material conductivity and thermal diffusivity. After the data is completed and standardized as described in step 2, it is input into the random forest laser absorption coefficient prediction model trained in step 4 to obtain the laser absorption coefficient required for the current laser processing numerical simulation.

[0017] Furthermore, the reasons for selecting the above parameters in step 1 are as follows:

[0018] The original data set includes: laser power, laser beam speed, spot radius, weld width, depth, electrical conductivity, thermal diffusivity, solid-liquid-gas three-phase temperature, etc. Factor analysis uses a small number of potential features with clear meanings to describe the relationship between features, grouping highly correlated features. This operation ensures that features within the same group are highly correlated, while features between different groups are less correlated. When partitioning the decision tree nodes, it avoids constructing decision trees with low accuracy, thereby improving the accuracy and convergence speed of the model. The formula is as follows:

[0019]

[0020] Among them, X i represents the i-th factor (feature), a ij is the factor loading matrix, F ij is the jth common factor, ε i is the special factor of the i-th factor; the common factor is obtained by maximum likelihood estimation; after factor analysis, the above parameters exclude several irrelevant parameters, and four characteristic groups with low inter-group correlation are obtained;

[0021] The laser absorption coefficient is strongly correlated with parameters such as laser and material properties. Laser power density, electrical conductivity, and thermal diffusivity are selected as important input parameters for predicting the laser absorption coefficient. The shape of the laser-processed shape is affected by the interaction between the laser and the material, and its shape coefficient can indirectly reflect the influence of the laser absorption coefficient. Therefore, laser energy density, shape coefficient, electrical conductivity, and thermal diffusivity are used as representative feature groups and important features for predicting the laser absorption coefficient. The definitions of laser energy density and shape coefficient are as follows:

[0022] Laser energy density:

[0023]

[0024] Where ED represents the laser energy density, P represents the laser power, v represents the laser beam speed, r represents the laser spot radius, and d represents the laser spot diameter. The above characteristics couple various process parameters in laser processing and can more accurately reflect their influence on the laser absorption coefficient.

[0025] Taking laser welding as an example, its forming coefficient is defined as follows:

[0026]

[0027] in, Indicates the weld forming coefficient, B is the weld width, and H is the weld depth; similar coefficients can be used in other areas of laser processing.

[0028] Furthermore, the steps of the raw data standardization process in step 2 are as follows:

[0029] S21: Calculate the mean μ and standard deviation σ of each variable respectively;

[0030] S22: Apply the normalization formula to normalize each data point. The formula is as follows:

[0031]

[0032] in, represents the i-th data point of the j-th variable, μ j is the mean of the jth variable, σ j is the standard deviation of the j-th variable;

[0033] S23: After standardization, the mean of each variable is 0 and the standard deviation is 1; the standardization operation is conducive to the rapid convergence of the subsequent mixed Gaussian model.

[0034] Furthermore, the steps for data expansion in step 3 are as follows:

[0035] S31: The standardized data mean is 0, the initial mean is simple to select, and the optimal number of clusters K is selected by the BIC criterion;

[0036] S32: Using a full diagonal covariance matrix, after standardizing the data, the elements in the covariance matrix become correlation coefficients, which helps the Gaussian mixture model capture the correlation between features more balanced, improves the model convergence speed, and reduces numerical calculation problems;

[0037] S33: Use the EM iterative method to optimize the model; the EM algorithm is divided into two steps: E step and M step, as follows:

[0038] The parameters of the Gaussian mixture model include Where K represents the number of clusters, π k 、μ k and ∑k represent the mixture weight, mean vector, and covariance matrix of the Kth Gaussian component. The EM algorithm, also known as the expectation maximization algorithm, optimizes the Gaussian mixture model by alternating two steps. The E step calculates the posterior responsibility of each sample for each Gaussian mixture component, and the M step updates the parameters based on the responsibility. The EM step is iterated until the increment of the log-likelihood is less than a threshold or the maximum number of iterations is reached. The EM algorithm ensures that the likelihood of a single point does not decrease and converges to a local optimum.

[0039] S34: Use the model to sample the original dataset to generate a new dataset, increasing the data volume by 10 times.

[0040] Furthermore, the training steps of the random forest laser absorption coefficient prediction model fused with factor analysis in step 4 are as follows:

[0041] S41: Use Bootstrap to randomly extract k sub-training sets from the training set, keeping the size consistent;

[0042] S42: Obtain the correlation between features through factor analysis method to form several feature groups;

[0043] S43: Selecting a number of input features from the feature group according to the number ratio to form a candidate feature subset;

[0044] S44: Use the optimal node decision function to obtain the best split feature and grow child nodes;

[0045] S45: recursively loop the above steps S42 to S44 until the maximum depth is reached, and integrate k decision trees to form a random forest model, where k is the number of decision trees in the random forest model. If k is 100, about 37% of the data is not selected, and this part of the data becomes the out-of-bag data, i.e., the test set;

[0046] S46: Use the mean squared error (MSE) of the 5-fold cross-validation model to determine whether the model meets the standard. If the accuracy meets the standard, the model training is successful. If the accuracy does not meet the requirements, grid search or random search is used to optimize the parameters. Since the training data is standardized and factor analyzed, the data quality is high, which can minimize or avoid parameter tuning and improve training efficiency.

[0047] Furthermore, in step 5, after the data is collected, it is supplemented and standardized as described in step 2, specifically using similar metal data to supplement the conductivity and thermal diffusivity, and using the training set mean and standard deviation to standardize the new data.

[0048] Another object of the present invention is to provide a random forest model based on factor analysis and an efficient prediction system for laser absorption coefficient, comprising:

[0049] A data set creation module is used to create a laser absorption coefficient data set using data from past experiments, simulations, and related literature. The data set includes laser energy density, form factor, electrical conductivity, thermal diffusivity, and laser absorption coefficient under different processes and welding materials.

[0050] The preprocessing module is used to perform preprocessing operations on the data set, including completion and standardization. The conductivity and thermal diffusivity of different materials in the data set may be missing, and similar metal data are used to complete them. The data set is also Z-score standardized to eliminate the differences between dimensions and facilitate subsequent data processing.

[0051] The incremental model construction module is used to build a data incremental model based on the Gaussian mixture model, generate a probability distribution of multiple Gaussian distribution linear mixtures to describe the original data, and randomly generate new samples through the distribution model to form an incremental data set;

[0052] The prediction model construction module is used to construct a random forest laser absorption coefficient prediction model based on fusion factor analysis. The laser absorption coefficient prediction model is trained using incremental data sets. Bootstrap is used to randomly extract k data sets with replacement. If k is 100, approximately 37% of the data is not selected. This data becomes the out-of-bag data, or test set. Fusion factor analysis is used to perform factor analysis on the incremental data sets to construct feature groups. After model training, an integrated model composed of multiple decision trees is obtained for predicting the laser absorption coefficient.

[0053] The calculation module is used to obtain the forming coefficient of the laser-processed product through experiments. At the same time, the laser energy density is calculated based on the laser power, movement speed and spot size of the laser used, and a set of data is formed in combination with the material conductivity and thermal diffusivity. After the data is completed and standardized as described in step 2, it is input into the random forest laser absorption coefficient prediction model trained in step 4 to obtain the laser absorption coefficient required for the current laser processing numerical simulation.

[0054] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for efficiently predicting the laser absorption coefficient based on the random forest model of factor analysis.

[0055] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the method for efficiently predicting the laser absorption coefficient based on the random forest model of factor analysis.

[0056] Another object of the present invention is to provide an information data processing terminal, which is used to implement the random forest model laser absorption coefficient efficient prediction system based on factor analysis.

[0057] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0058] This paper provides a random forest laser absorption coefficient prediction method that integrates factor analysis. This method comprehensively considers the influence of factors such as laser type and material properties, and draws on past experimental and simulation experience to establish a laser absorption coefficient dataset. This dataset is expanded using a mixed Gaussian model, which is then used to train a random forest regression prediction model. Factor analysis is used to optimize feature selection, improving training efficiency and model interpretability. This method achieves efficient and meaningful laser absorption coefficient prediction, providing a highly effective parameter estimation method for laser processing numerical simulation technology.

[0059] In this method, a random forest laser absorption coefficient prediction method that integrates factor analysis is used. The laser absorption coefficient dataset includes four parameters: laser power density, forming coefficient, electrical conductivity, and thermal diffusivity. It comprehensively considers various factors affecting the laser absorption coefficient in laser processing, reduces the dimension of the dataset, and has low feature correlation between the data, providing a high-quality training set for subsequent random forest model training.

[0060] In this method, a random forest laser absorption coefficient prediction method that integrates factor analysis is used. A Gaussian mixture model is used to generate a probability distribution of multiple Gaussian distribution linear mixtures to describe the original data, and random sampling is used to generate new samples to form an incremental data set, thereby improving the training effect of the random forest model and avoiding overfitting. Through factor analysis, features with strong correlation are grouped together to avoid the generation of weak decision trees, improve the accuracy and convergence speed of the model, and achieve an efficient model training method.

[0061] This method provides a feasible connection method for laser processing numerical twins. In the field of laser processing, the connection between virtual and physical space is difficult and complex due to the complex interaction between lasers and materials. This method provides a feasible way to predict the energy distribution in virtual space from physical space data. By combining physical space data such as laser energy density and conductivity with the random forest laser absorption coefficient using fusion factor analysis to predict the laser absorption coefficient required for numerical simulation, combined with a ray tracing model, the energy distribution of the laser in physical space can be restored.

[0062] This invention fills a technological gap in the accurate prediction of the laser absorption coefficient in traditional laser welding numerical simulations. Specifically, a random forest prediction model based on factor analysis is proposed for different welding processes and material properties. In terms of feature selection, the influence of process parameters, material properties, and welding results on the laser absorption coefficient is comprehensively considered. Factor analysis is used to group features, and a random forest model is employed to achieve efficient prediction of the laser absorption coefficient for laser welding.

[0063] This invention addresses the technical bias in the empirical selection of laser absorption coefficients in traditional laser welding numerical simulations. Previous studies have primarily relied on empirical selection of laser absorption coefficients, which has led researchers to constantly adjust the absorption coefficient based on simulation results, reducing research efficiency. The present invention leverages the experience gained from previous studies to construct a model training dataset. Laser energy density, forming coefficient, electrical conductivity, and thermal diffusivity are selected as model feature inputs, and factor analysis is used to avoid selecting features with high feature correlation, thereby improving model training efficiency. Finally, a random forest regression model is used to achieve efficient prediction of the laser absorption coefficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 This is a flow chart of an efficient prediction method for laser absorption coefficient using a random forest model based on factor analysis provided by an embodiment of the present invention.

[0065] Figure 2 This is a structural block diagram of a system for efficiently predicting laser absorption coefficient using a random forest model based on factor analysis provided by an embodiment of the present invention.

[0066] Figure 3 Schematic diagram of the experimental equipment and numerical model provided by the embodiment of the present invention.

[0067] Figure 4 Schematic diagram of the random forest model architecture for fusion factor analysis provided by an embodiment of the present invention.

[0068] Figure 5 This is a comparison chart of prediction results of a certain example provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0069] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0070] like Figure 1 As shown, an embodiment of the present invention provides an efficient prediction method for laser absorption coefficient based on a random forest model of factor analysis, comprising the following steps:

[0071] S101, establishing a laser absorption coefficient dataset using data from past experiments, simulations, and relevant literature; the dataset includes laser energy density, form factor, electrical conductivity, thermal diffusivity, and laser absorption coefficient under different processes and welding materials;

[0072] S102, preprocessing the data set, including completion and standardization. The data set may contain missing values ​​for the conductivity and thermal diffusivity of different materials, which are then completed using data from similar metals. The data set is also Z-score standardized to eliminate dimensional differences and facilitate subsequent data processing.

[0073] S103, constructing a data increment model based on a Gaussian mixture model, generating a probability distribution of a linear mixture of multiple Gaussian distributions to describe the original data, and randomly generating new samples to form an incremental data set through the distribution model;

[0074] S104, constructing a random forest laser absorption coefficient prediction model based on fusion factor analysis, and using the incremental data set to train the laser absorption coefficient prediction model; using Bootstrap with replacement to randomly extract k data sets, if k is 100, approximately 37% of the data is not selected, and this part of the data becomes the out-of-bag data, i.e., the test set; using fusion factor analysis to perform factor analysis on the incremental data set to construct a feature group, and after model training, an integrated model consisting of multiple decision trees is obtained for predicting the laser absorption coefficient;

[0075] In step S105, the forming coefficient of the laser-processed product is obtained through experiments. At the same time, the laser energy density is calculated based on the laser power, moving speed, and spot size of the laser used, and a set of data is formed in combination with the material conductivity and thermal diffusivity. After the data is completed and standardized as described in step S102, it is input into the random forest laser absorption coefficient prediction model trained in step S104 to obtain the laser absorption coefficient required for the current laser processing numerical simulation.

[0076] The reasons for selecting the above parameters in S101 provided in the embodiment of the present invention are as follows:

[0077] The laser absorption coefficient is strongly correlated with parameters such as laser and material properties. Laser power density, electrical conductivity, and thermal diffusivity are selected as important input parameters for predicting the laser absorption coefficient. The shape of the laser-processed shape is affected by the interaction between the laser and the material, and its shape coefficient can indirectly reflect the influence of the laser absorption coefficient. Therefore, laser energy density, shape coefficient, electrical conductivity, and thermal diffusivity are used as important characteristics for predicting the laser absorption coefficient. The definitions of laser energy density and shape coefficient are as follows:

[0078] Laser energy density:

[0079]

[0080] Where ED represents the laser energy density, P represents the laser power, v represents the laser beam speed, r represents the laser spot radius, and d represents the laser spot diameter. The above characteristics couple various process parameters in laser processing and can more accurately reflect their influence on the laser absorption coefficient.

[0081] Taking laser welding as an example, its forming coefficient is defined as follows:

[0082]

[0083] in, Indicates the weld forming coefficient, B is the weld width, and H is the weld depth; similar coefficients can be used in other areas of laser processing.

[0084] The steps of the raw data normalization process in S102 provided in the embodiment of the present invention are as follows:

[0085] S21: Calculate the mean μ and standard deviation σ of each variable respectively;

[0086] S22: Apply the normalization formula to normalize each data point. The formula is as follows:

[0087]

[0088] in, represents the i-th data point of the j-th variable, μ j is the mean of the jth variable, σ j is the standard deviation of the j-th variable;

[0089] S23: After standardization, the mean of each variable is 0 and the standard deviation is 1; the standardization operation is conducive to the rapid convergence of the subsequent mixed Gaussian model.

[0090] The steps of data expansion in S103 provided in the embodiment of the present invention are as follows:

[0091] S31: The standardized data mean is 0, the initial mean is simple to select, and the optimal number of clusters K is selected by the BIC criterion;

[0092] S32: Using a full diagonal covariance matrix, after standardizing the data, the elements in the covariance matrix become correlation coefficients, which helps the Gaussian mixture model capture the correlation between features more balanced, improves the model convergence speed, and reduces numerical calculation problems;

[0093] S33: Use the EM iterative method to optimize the model; the EM algorithm is divided into two steps: E step and M step, as follows:

[0094] The parameters of the Gaussian mixture model include Where K represents the number of clusters, π k 、μ kand ∑k represent the mixture weight, mean vector, and covariance matrix of the Kth Gaussian component. The EM algorithm, also known as the expectation maximization algorithm, optimizes the Gaussian mixture model by alternating two steps. The E step calculates the posterior responsibility of each sample for each Gaussian mixture component, and the M step updates the parameters based on the responsibility. The EM step is iterated until the increment of the log-likelihood is less than a threshold or the maximum number of iterations is reached. The EM algorithm ensures that the likelihood of a single point does not decrease and converges to a local optimum.

[0095] S34: Use the model to sample the original dataset to generate a new dataset, increasing the data volume by 10 times.

[0096] The training steps of the random forest laser absorption coefficient prediction model for fusion factor analysis in S104 provided by the embodiment of the present invention are as follows:

[0097] S41: Use Bootstrap to randomly extract k sub-training sets from the training set, keeping the size consistent;

[0098] S42: Obtain the correlation between features through factor analysis method to form several feature groups;

[0099] S43: Selecting a number of input features from the feature group according to the number ratio to form a candidate feature subset;

[0100] S44: Use the optimal node decision function to obtain the best split feature and grow child nodes;

[0101] S45: recursively loop the above steps S42 to S44 until the maximum depth is reached, and integrate k decision trees to form a random forest model, where k is the number of decision trees in the random forest model. If k is 100, about 37% of the data is not selected, and this part of the data becomes the out-of-bag data, i.e., the test set;

[0102] S46: Use the mean squared error (MSE) of the 5-fold cross-validation model to determine whether the model meets the standard. If the accuracy meets the standard, the model training is successful. If the accuracy does not meet the requirements, grid search or random search is used to optimize the parameters. Since the training data is standardized and factor analyzed, the data quality is high, which can minimize or avoid parameter tuning and improve training efficiency.

[0103] The effects of the factor analysis described in S104 are as follows:

[0104] Factor analysis uses a small number of potential features with clear meanings to describe the relationship between features, grouping highly correlated features into a group. This ensures that features within the same group are highly correlated, while features in different groups are less correlated. This avoids constructing decision trees with low accuracy when dividing decision tree nodes, thereby improving the accuracy and convergence speed of the model. The formula is as follows:

[0105]

[0106] Among them, X i represents the i-th factor (feature), a ij is the factor loading matrix, F ij is the jth common factor, ε i is the special factor of the i-th factor; the common factors are obtained by maximum likelihood estimation.

[0107] In S105 provided by the embodiment of the present invention, after the data is collected, it is supplemented and standardized as described in S102, specifically using similar metal data to supplement the conductivity and thermal diffusivity, and using the training set mean and standard deviation to standardize the new data.

[0108] like Figure 2 As shown, an embodiment of the present invention provides an efficient prediction system for laser absorption coefficient based on a random forest model of factor analysis, including:

[0109] A data set creation module is used to create a laser absorption coefficient data set using data from past experiments, simulations, and related literature. The data set includes laser energy density, form factor, electrical conductivity, thermal diffusivity, and laser absorption coefficient under different processes and welding materials.

[0110] The preprocessing module is used to perform preprocessing operations on the data set, including completion and standardization. The conductivity and thermal diffusivity of different materials in the data set may be missing, and similar metal data are used to complete them. The data set is also Z-score standardized to eliminate the differences between dimensions and facilitate subsequent data processing.

[0111] The incremental model construction module is used to build a data incremental model based on the Gaussian mixture model, generate a probability distribution of multiple Gaussian distribution linear mixtures to describe the original data, and randomly generate new samples through the distribution model to form an incremental data set;

[0112] The prediction model construction module is used to construct a random forest laser absorption coefficient prediction model based on fusion factor analysis. The laser absorption coefficient prediction model is trained using incremental data sets. Bootstrap is used to randomly extract k data sets with replacement. If k is 100, approximately 37% of the data is not selected. This data becomes the out-of-bag data, or test set. Fusion factor analysis is used to perform factor analysis on the incremental data sets to construct feature groups. After model training, an integrated model composed of multiple decision trees is obtained for predicting the laser absorption coefficient.

[0113] The calculation module is used to obtain the forming coefficient of the laser-processed product through experiments, and at the same time calculate the laser energy density based on the laser power, moving speed and spot size of the laser used, and form a set of data in combination with the material conductivity and thermal diffusivity; after the data is completed and standardized as described in S102, it is input into the random forest laser absorption coefficient prediction model trained in S104 to obtain the laser absorption coefficient required for the current laser processing numerical simulation.

[0114] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for efficiently predicting the laser absorption coefficient based on the random forest model of factor analysis.

[0115] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to perform the steps of the method for efficiently predicting the laser absorption coefficient based on the random forest model of factor analysis.

[0116] Another object of the present invention is to provide an information data processing terminal, which is used to implement the random forest model laser absorption coefficient efficient prediction system based on factor analysis.

[0117] The present invention is specifically implemented:

[0118] The experimental equipment and numerical model involved in this embodiment are as follows Figure 3 As shown in the attached figure, the random forest laser absorption coefficient prediction model architecture of fusion factor analysis is shown in the attached figure. Figure 4 shown.

[0119] As attached Figure 3 The experimental equipment shown in the figure is used in this case to study the laser welding process. The power range of the welding equipment is: 1-5KW, the welding speed range is: 5-50mm / s, and the spot diameter is 500-1000um. The materials selected are steel, aluminum, titanium, and copper. The electrical conductivity and thermal diffusivity of the relevant materials are shown in the following table:

[0120] Material Conductivity%IACS <![CDATA[Thermal diffusivity mm 2 / s]]> Material Conductivity%IACS <![CDATA[Thermal diffusivity mm 2 / s <!-- 8 -->]]> mild steel 7.0 15 2024 aluminum alloy 35.0 55.0 304 stainless steel 2.0 3.8 6061 aluminum alloy 45.0 70.0 410 stainless steel 3.0 5.0 TC4 titanium alloy 1.3 3.0 2195 aluminum alloy 22.0 45.0 TA15 titanium alloy 2.0 3.0 2219 aluminum alloy 32.0 50.0 brass 30 40

[0121] Many studies have been conducted on the laser absorption coefficients of the above materials under different processes. Calorimetry can obtain the average absorption rate of the material in the high temperature zone. Figure 3As shown, numerical simulations can also yield relatively accurate absorption rates by comparing weld profiles. In summary, this method, through numerical simulations, physical experiments, and literature review, determined 300 sets of laser absorption coefficients for different processes and materials. Furthermore, weld penetration, weld width, and weld profile are closely related to the energy input during welding. This method considers these as important characteristics of the laser absorption coefficient and incorporates them into the dataset as input variables. This yielded 300 datasets of laser absorption coefficients with four variables.

[0122] Missing items in the dataset were completed. The electrical conductivity and thermal diffusivity of some materials were missing. Since similar metals have similar physical properties, the electrical conductivity and thermal diffusivity of similar metals were used to supplement the missing items. The dataset was normalized using Z-score to eliminate differences in dimensions and facilitate subsequent data processing.

[0123] The optimal number of clusters, K, was selected using the BIC criterion. In this case, the number of clusters was 4, and the initial weight of each Gaussian component was 1 / K. After inputting the standardized data, the covariance matrix of the overall data was calculated. A full diagonal covariance matrix was used to improve the model's convergence speed. The model was optimized using the EM iterative method. In this case, the maximum increment of the log-likelihood was set to 10-3, and the maximum number of iterations was set to 100 to prevent infinite loops. After the model was completed, a probability distribution of multiple linear mixtures of Gaussian distributions was generated to describe the original data, and new samples were randomly sampled to form an incremental data set. This increased the data volume by 10 times, resulting in a sample size of 3,000. This allowed for effective training of the random forest model, improving the training effect of the prediction model and avoiding overfitting.

[0124] Figure 5 The random forest laser absorption coefficient prediction model, incorporating factor analysis, uses factor analysis to determine feature correlations. This grouping of strongly correlated features avoids constructing inaccurate decision trees during node splitting, thereby improving model accuracy and convergence speed. During model training, 100 sub-training sets were randomly sampled with replacement from the training set, maintaining the same size as the original training set. Approximately 37% of the original data set was excluded as out-of-bag data and the test set. These 100 sub-training sets were used to train 100 decision trees to form the random forest model. When splitting the decision tree nodes, feature inputs were proportionally selected from the feature groups. The optimal node decision function was used to calculate the mean squared error (MSE) to determine the optimal splitting feature and grow the child nodes. This step was recursively repeated until the maximum depth was reached, at which point the 100 decision trees were ensembled to form the random forest model. The mean squared error (MSE) from a 5-fold cross-validation validation was used to determine model compliance. If the accuracy met the target, the model was considered trained successfully. If not, grid search or random search was used to optimize the parameters. The following table lists the parameters for the random forest model:

[0125] parameter Value k (number of trees) 100 d (maximum depth) 10 Split (minimum number of samples required for node splitting) 5 Leaf (minimum number of samples for leaf nodes) 2 Max_features (number of features selected by tree splitting) Random Decrease

[0126] A set of laser welding experiments were designed, with a laser power of 4KW, a moving speed of 40mm / s, a spot diameter of 500um, an energy density of 509.3 (J / mm), and a material of TC4 titanium alloy with electrical conductivity and thermal diffusivity of 1.3 and 3.0. The weld formation coefficient obtained from the welding experiment was 0.6. The laser absorption coefficient obtained by the model calculation was 0.187. Figure 3 The numerical model in the paper obtained a laser absorption coefficient of 0.19, which proved that the prediction model was accurate and had certain practical value, greatly shortening the model debugging time.

[0127] The present invention has two application prospects. First, based on the technical difficulties of empirical selection of laser absorption coefficient, the present invention proposes a prediction method based on random forest model. In actual application scenarios, researchers can use this method to obtain a more accurate laser absorption coefficient, reduce model adjustment time and improve research efficiency. Secondly, in terms of data processing, factor analysis is used to use a small number of potential features with clear meanings to describe the relationship between features, and the features with strong correlation are divided into a group. This gives the present invention a flexibility that can be naturally grafted into the prediction of other model parameters. The numerical model of laser welding has a large number of complex parameters that change dynamically, such as: nonlinear physical parameters such as viscosity and surface tension coefficient, phase interface thickness, convective heat transfer coefficient, radiation emissivity, etc. In summary, the present invention has certain application value and can be applied to the prediction of similar parameters.

[0128] The main body of the present invention is based on three methods: factor analysis, Gaussian mixture model and random forest regression model. First, factor analysis is a feature dimension compression method for data, which combines highly correlated features into one factor. For example, in the present invention, the highly correlated laser power, beam moving speed, and spot size are combined into laser energy density, thereby achieving data set dimensionality reduction. This method has applications in many fields and the technology is relatively mature. Secondly, the Gaussian mixture model is used to achieve data increment, which is common in other fields and the technology is relatively mature. Finally, the random forest regression model is also a relatively mature machine learning method. This model has a low demand for data sample size. After factor analysis and data expansion, this method can achieve efficient training of the model. Appendix Figure 5 A set of comparison graphs is provided. The numerical simulation results, using the laser absorption coefficient predicted by this method, show that the weld penetration and formed morphology are essentially consistent. In summary, from the perspective of technical feasibility, the three methods described above are relatively mature methods in other fields, with relatively low technical difficulty, and can achieve the technical effects described in this invention. In terms of actual results, the model prediction results are consistent with the actual welding results, demonstrating that the prediction results of this invention are relatively accurate and can achieve accurate prediction of laser welding welds.

[0129] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0130] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.

Claims

1. A laser absorption coefficient prediction method based on factor analysis and random forest, characterized in that: The following steps are involved: Create a data set including laser energy density, shape factor, electrical conductivity, thermal diffusivity and laser absorption coefficient; Perform missing value imputation on the dataset and normalize each feature to a standardized form with a mean of zero and a variance of one; generating incremental data based on a Gaussian mixture model and merging the incremental data with the data set to obtain an extended data set; Performing factor analysis on the extended data set to construct a feature group, and training a random forest through bootstrap sampling to obtain a prediction model; The laser processing parameters and material properties are collected and input into the prediction model after being processed in step 2 to output the laser absorption coefficient.

2. The method according to claim 1, characterized in that In step three, the Bayesian information criterion is used to select the number of clusters, and the model parameters are iterated through the expectation maximization algorithm until the log-likelihood increment is lower than the preset threshold.

3. The method according to claim 1, characterized in that The random forest trained in step 4 contains a preset number of decision trees. The out-of-bag data is used to evaluate the model error, and the training ends when the mean square error meets the preset threshold.

4. A laser absorption coefficient prediction system based on factor analysis and random forest, characterized in that: It includes a data set establishment unit, a preprocessing unit, an incremental data generation unit, a model training unit and a prediction unit; each unit cooperates to perform the steps of claim 1 to output the laser absorption coefficient.

5. The system according to claim 4, characterized in that The incremental data generation unit is configured to implement data expansion in the manner described in claim 2.

6. A laser absorption coefficient data incremental modeling method, characterized in that: The following steps are involved: Perform standardization on raw data; The number of clusters in the Gaussian mixture model was determined based on the Bayesian Information Criterion; The expectation maximization algorithm is used to estimate the model parameters; Synthetic samples are generated based on the obtained model until the amount of data reaches a predetermined multiple of the original dataset capacity.

7. A laser absorption coefficient prediction model training method, characterized in that: The following steps are involved: Factor analysis was performed on the extended data to form feature groups; Perform bootstrap sampling on the training set to generate multiple sub-training sets; Train a decision tree on each sub-training set and ensemble it into a random forest; The performance is evaluated using out-of-bag data and training is completed when a preset error threshold is met.

8. A computing device comprising a memory and a processor, characterized in that: The memory stores instructions that, when executed on the processor, cause the computing device to implement the method of claim 1 . 9 . A computer-readable storage medium having a computer program recorded thereon, wherein when the program is executed by a processor, the processor is caused to execute the method according to claim 1 .

10. A prediction server, characterized in that: The method comprises the system according to claim 4 and a network interface, wherein the network interface is configured to receive laser processing parameters and material properties and output the laser absorption coefficient generated by the system.

Citation Information

Cited By

  • A method for laser machining parameter optimization for tool surface texturing

    CN122389001A