Method, apparatus and electronic device for constructing alloy data sample set

By performing data cleaning, feature screening, feature addition and feature standardization on the initial alloy data sample set, the target alloy data sample set that meets the needs of the multi-dimensional feature alloy model is constructed, solving the complex nonlinear problem between alloy preparation strategy and performance, and improving the prediction accuracy of the machine learning model.

CN115274022BActive Publication Date: 2025-06-10SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211048392.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-06-10
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively build an alloy data sample set that meets the needs of multi-dimensional characteristic alloy models, and cannot effectively solve the complex nonlinear problem between alloy preparation strategies and performance.

Method used

By obtaining the initial alloy data sample set, data cleaning, feature screening and feature addition are carried out, and feature standardization is carried out to construct the target alloy data sample set. The method includes screening alloy casting process data from the literature, extracting alloy elements, heat treatment parameters and performance data, performing data cleaning and feature screening, adding thermodynamic features that affect the mechanical properties of the alloy, and determining the importance of features through machine learning algorithms, and finally performing feature standardization.

Benefits of technology

The built target alloy data sample set can meet the needs of multidimensional characteristic alloy models, help solve the complex nonlinear problems between alloy preparation strategies and performance, and improve the prediction accuracy of machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115274022B_ABST
    Figure CN115274022B_ABST
Patent Text Reader

Abstract

The present application provides a method, an apparatus and an electronic device for constructing an alloy data sample set. The method includes: obtaining an initial alloy data sample set; the samples in the initial alloy data sample set include alloy preparation strategies and alloy properties corresponding to the alloy preparation strategies; wherein, the alloy preparation strategies include: alloy element contents and heat treatment parameters; cleaning the data of the samples in the initial alloy data sample set to obtain a first intermediate data sample set; performing feature screening and feature addition on the samples in the first intermediate data sample set to obtain a second intermediate data sample set; performing feature standardization processing on the samples in the second intermediate data sample set to obtain a target alloy data sample set corresponding to the initial alloy data sample set. The present application can meet the construction requirements of a multi-dimensional feature alloy model and is beneficial to solving the complex non-linear problem between alloy preparation strategies and properties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of materials, and in particular, to a method, apparatus, and electronic device for constructing an alloy data sample set. Background Art

[0002] With the progress of the times and the development of aviation technology, modern industry has put forward higher requirements for the performance of alloys. However, there are many alloying elements and complex and diverse strengthening mechanisms. This makes the traditional experimental trial-and-error method time-consuming and laborious. How to discover new high-performance alloys remains a major problem. Therefore, computational methods play an increasingly important role in the research of alloy materials. Compared with traditional computational methods, machine learning methods have their special advantages, can quickly solve complex non-linear relationships, and can easily update the machine learning model by replacing relevant data sets.

[0003] A very crucial point in machine learning modeling is that the construction of the data set and the quality of the features usually determine the upper limit of a machine learning model, and the relevant prediction algorithms of machine learning only try to approach this upper limit as much as possible. This shows that the construction of the data set is a key factor affecting the accuracy of the machine learning model.

[0004] Due to the complex non-linear relationship between alloy composition and mechanical properties, multi-feature modeling of alloys by machine learning is still a key issue. Among them, the reasonable construction of the data set is of utmost importance and is the basis for determining the accuracy of the machine learning model. In addition, the performance of high-strength aluminum alloys is always affected by the heat treatment system. Therefore, formulating a suitable feature extraction strategy also plays a crucial role in constructing a reasonable machine learning model. However, the current application of machine learning modeling for aluminum alloys mainly focuses on the prediction of alloy composition and the study of single performance. Therefore, the data features of the constructed sample set are relatively single, unable to meet the needs of the multi-dimensional feature alloy model, and not conducive to solving the complex non-linear problem between alloy preparation strategies and performance. Summary of the Invention

[0005] The purpose of the present application is to provide a method, apparatus, and electronic device for constructing an alloy data sample set, which can meet the construction requirements of a multi-dimensional feature alloy model and is conducive to solving the complex non-linear problem between alloy preparation strategies and performance.

[0006] In a first aspect, an embodiment of the present application provides a method for constructing an alloy data sample set, the method including: obtaining an initial alloy data sample set; samples in the initial alloy data sample set include alloy preparation strategies and alloy properties corresponding to the alloy preparation strategies; wherein, the alloy preparation strategies include: alloy element contents and heat treatment parameters; cleaning the data of the samples in the initial alloy data sample set to obtain a first intermediate data sample set; performing feature screening and feature addition on the samples in the first intermediate data sample set to obtain a second intermediate data sample set; performing feature standardization processing on the samples in the second intermediate data sample set to obtain a target alloy data sample set corresponding to the initial alloy data sample set.

[0007] In a preferred embodiment of the present application, the step of obtaining the initial alloy data sample set includes: screening target documents describing alloy casting processes from documents on aluminum alloys; screening data containing the corresponding relationship between alloy preparation strategies and alloy properties from the target documents; establishing an initial alloy data sample set based on the data containing the corresponding relationship between alloy preparation strategies and alloy properties.

[0008] In a preferred embodiment of the present application, the step of cleaning the data of the samples in the initial alloy data sample set to obtain a first intermediate data sample set includes: extracting alloy elements with the occurrence times reaching the times threshold and the content reaching the content threshold from the initial alloy data sample set, extracting heat treatment parameters with both solution treatment and aging, and extracting alloy properties including at least one of tensile strength, yield strength, and elongation; generating a first intermediate data sample set based on the extracted alloy elements and corresponding contents, heat treatment parameters, and alloy properties.

[0009] In a preferred embodiment of the present application, the step of performing feature screening and feature addition on the samples in the first intermediate data sample set to obtain a second intermediate data sample set includes: obtaining multiple thermodynamic features affecting the mechanical properties of alloys; determining importance parameters corresponding to the multiple thermodynamic features respectively based on a machine learning algorithm; determining the thermodynamic feature corresponding to the largest importance parameter as the target thermodynamic feature; for each sample, determining the feature value of the sample under the target thermodynamic feature and adding the feature value to the sample to obtain a second intermediate data sample set.

[0010] In a preferred embodiment of the present application, the multiple thermodynamic features include: mixing enthalpy, mixing entropy, electronegativity, lattice distortion energy, valence electron concentration; the target thermodynamic feature is the atomic size difference feature.

[0011] In a preferred embodiment of the present application, the step of determining the feature value of the sample under the target thermodynamic feature includes: calculating the feature value of the sample under the target thermodynamic feature according to the following formula:

[0012]

[0013] Among them, δr represents the eigenvalue of the sample under the atomic size difference feature; c i and c j are the atomic percentages corresponding to the i-th element and the j-th element in the sample respectively, r i represents the atomic radius of the i-th element; n represents the number of elements in the sample; represents the average atomic radius of the n elements in the sample.

[0014] In a preferred embodiment of the present application, the step of performing feature standardization processing on the samples in the second intermediate data sample set includes: performing feature standardization processing on the samples in the second intermediate data sample set according to the following formula:

[0015] i = 1, 2, …, M; j = 1, 2, …, N;

[0016] Among them, N represents the number of samples; M represents the number of features in the sample; represents the feature after standardization processing of the i-th feature in the j-th sample; represents the initial feature of the i-th feature in the j-th sample; μ (i) represents the average value corresponding to the i-th feature among N samples; σ (i) represents the standard deviation corresponding to the i-th feature among N samples.

[0017] In a second aspect, an apparatus for constructing an alloy data sample set provided by an embodiment of the present application includes: an initial sample set acquisition module, configured to acquire an initial alloy data sample set; the samples in the initial alloy data sample set include alloy preparation strategies and alloy properties corresponding to the alloy preparation strategies; among them, the alloy preparation strategies include: alloy element content and heat treatment parameters; a data cleaning module, configured to perform data cleaning on the samples in the initial alloy data sample set to obtain a first intermediate data sample set; a feature screening and adding module, configured to perform feature screening and feature adding on the samples in the first intermediate data sample set to obtain a second intermediate data sample set; a feature standardization module, configured to perform feature standardization processing on the samples in the second intermediate data sample set to obtain a target alloy data sample set corresponding to the initial alloy data sample set.

[0018] In a third aspect, an embodiment of the present application further provides an electronic device, including a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method described in the first aspect above.

[0019] Fourthly, an embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the method described in the first aspect above.

[0020] In the method, device, and electronic device for constructing an alloy data sample set provided by the embodiments of the present application, an initial alloy data sample set is first obtained; the samples in the initial alloy data sample set include alloy preparation strategies and the alloy properties corresponding to the alloy preparation strategies; among them, the alloy preparation strategies include: alloy element content and heat treatment parameters; then the samples in the initial alloy data sample set are subjected to data cleaning to obtain a first intermediate data sample set; then the samples in the first intermediate data sample set are subjected to feature screening and feature addition to obtain a second intermediate data sample set; finally, the samples in the second intermediate data sample set are subjected to feature standardization processing to obtain a target alloy data sample set corresponding to the initial alloy data sample set. In the embodiments of the present application, by performing data cleaning, feature screening, feature addition, and feature standardization processing on the initial alloy data sample set, a target alloy data sample set that can meet the construction requirements of a multi-dimensional feature alloy model is obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of a method for constructing an alloy data sample set provided by an embodiment of the present application;

[0023] Figure 2 It is a schematic diagram of a process for constructing an alloy data sample set provided by an embodiment of the present application;

[0024] Figure 3 It is a flowchart of a process for obtaining an initial alloy data sample set in a method for constructing an alloy data sample set provided by an embodiment of the present application;

[0025] Figure 4 It is a flowchart of a process for feature screening and addition in a method for constructing an alloy data sample set provided by an embodiment of the present application;

[0026] Figure 5 It is a structural block diagram of a device for constructing an alloy data sample set provided by an embodiment of the present application;

[0027] Figure 6 Schematic diagram of a structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0028] The technical solutions of the present application will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0029] Currently, the application of machine learning modeling for aluminum alloys mainly focuses on the prediction of alloy compositions and the study of single properties. Therefore, the data features of the constructed sample set are relatively single, unable to meet the requirements of the multi-dimensional feature alloy model, and not conducive to solving the complex non-linear problems between alloy preparation strategies and properties. Based on this, the embodiments of the present application provide a method, device, and electronic device for constructing an alloy data sample set, which can meet the construction requirements of the multi-dimensional feature alloy model and is conducive to solving the complex non-linear problems between alloy preparation strategies and properties.

[0030] To facilitate the understanding of this embodiment, a method for constructing an alloy data sample set disclosed in the embodiments of the present application will be introduced in detail first.

[0031] Figure 1 Flowchart of a method for constructing an alloy data sample set provided by an embodiment of the present application. The method specifically includes the following steps:

[0032] Step S102, obtain an initial alloy data sample set; the samples in the initial alloy data sample set include alloy preparation strategies and the alloy properties corresponding to the alloy preparation strategies; wherein, the alloy preparation strategies include: alloy element content and heat treatment parameters.

[0033] See Figure 2 As shown, the alloy elements include: 35 kinds such as Al, Zn, Mg, and Gu. The heat treatment parameters, that is, the heat treatment system, include: solution treatment and aging treatment. The alloy properties include: 10 kinds such as tensile strength, yield strength, elongation, hardness, and conductivity.

[0034] Step S104, perform data cleaning on the samples in the initial alloy data sample set to obtain a first intermediate data sample set.

[0035] For example, select alloy elements with a data volume greater than 100 and an average content greater than 0.01 wt%; heat treatment parameters that must include both solution treatment and aging treatment; the performance targets that are most concerned about in this study.

[0036] Step S106: Screen and add features to the samples in the first intermediate data sample set to obtain the second intermediate data sample set.

[0037] In specific implementation, relevant physical features are added according to the strengthening mechanism, the importance of features is calculated using machine learning algorithms, and then combined with materials science knowledge to judge the feature calculation results and add features. The above-mentioned multiple thermodynamic features may include: mixing enthalpy, mixing entropy, electronegativity, lattice distortion energy, valence electron concentration, atomic size difference. Through screening, the feature values under the atomic size difference feature can finally be added to the samples.

[0038] Step S108: Perform feature standardization processing on the samples in the second intermediate data sample set to obtain the target alloy data sample set corresponding to the initial alloy data sample set.

[0039] The features involved in this embodiment include composition features, heat treatment features, and material property features, and standardization processing is performed on each feature.

[0040] In the method for constructing the alloy data sample set provided by the embodiment of the present application, through data cleaning, feature screening, feature addition, and feature standardization processing on the initial alloy data sample set, a target alloy data sample set that can meet the construction requirements of the multi-dimensional feature alloy model is obtained. In the embodiment of the present application, aluminum alloy is used as the research object, a reasonable data set construction method applicable to machine learning modeling is proposed, and the screening of key features affecting alloy performance is completed, thus laying a good foundation for the reasonable design of the machine learning model. The data set construction method proposed in this solution helps to solve the complex non-linear problem between alloy preparation strategies and performance, and can play a significant role in improving the prediction accuracy of the machine learning model.

[0041] The embodiment of the present application also provides another method for constructing an alloy data sample set, which is implemented on the basis of the above embodiment; this embodiment focuses on describing the processes of data cleaning, feature screening, feature addition, and feature standardization.

[0042] See Figure 3 As shown, the steps for obtaining the initial alloy data sample set include:

[0043] Step S302: Screen target documents describing alloy casting processes from the literature on aluminum alloys;

[0044] Step S304: Screen data containing the corresponding relationship between alloy preparation strategies and alloy performance from the target documents;

[0045] Step S306: Based on the data containing the corresponding relationship between alloy preparation strategies and alloy performance, establish an initial alloy data sample set.

[0046] In specific implementation, for hundreds of thousands of scientific reports on aluminum alloys that can be accessed, only the literature with the preparation method unified as the casting process was investigated. At the same time, it was required that detailed data on element content, heat treatment system, and mechanical properties must be provided synchronously in the article. On this basis, a data set consisting of three parts: element content, heat treatment parameters, and material properties was established, that is, the initial alloy data sample set.

[0047] The steps of cleaning the samples in the initial alloy data sample set to obtain the first intermediate data sample set include: extracting alloy elements that appear more than the frequency threshold and have a content reaching the content threshold from the initial alloy data sample set, extracting heat treatment parameters with both solution treatment and aging, and extracting alloy properties including at least one of tensile strength, yield strength, and elongation; generating the first intermediate data sample set based on the extracted alloy elements and their corresponding contents, heat treatment parameters, and alloy properties.

[0048] In specific implementation, only the elements that appear more than 100 times in the data set and have an average content greater than 0.01 wt% are selected as the final composition characteristics. In addition, considering the importance of the heat treatment system, solution treatment and aging are selected as the heat treatment characteristics in the embodiments of the present application. Finally, the performance targets that this embodiment focuses on the most are tensile strength, yield strength, and elongation.

[0049] See Figure 4 As shown, the steps of performing feature screening and feature addition on the samples in the first intermediate data sample set to obtain the second intermediate data sample set include the following steps:

[0050] Step S402, obtaining multiple thermodynamic characteristics that affect the mechanical properties of the alloy; the multiple thermodynamic characteristics include: mixing enthalpy, mixing entropy, electronegativity, lattice distortion energy, valence electron concentration, atomic size difference;

[0051] Step S404, determining the importance parameters corresponding to the multiple thermodynamic characteristics based on a machine learning algorithm. In specific implementation, two algorithms, MLP and SVR, can be used to calculate the feature importance respectively, and according to the calculation results, determine the key features that most significantly affect the alloy properties.

[0052] Step S406, determining the thermodynamic characteristic corresponding to the largest importance parameter as the target thermodynamic characteristic.

[0053] Thermodynamic feature screening and addition are the most crucial steps in the embodiments of this application to improve the prediction accuracy of machine learning models. Taking aluminum alloy as an example, solid solution strengthening is one of the key factors affecting the mechanical properties of Al-Zn-Mg-Cu series alloys. Based on the solubility theory proposed by Hume-Rothery, the difference in the atomic sizes of alloying elements has a significant impact on solubility. A significant difference in atomic size will cause severe lattice distortion, thus sharply increasing the energy of the alloy system. The high energy will drive the unstable or metastable solid solution to transform into the thermodynamically stable intermetallic compound. In addition, a larger atomic size difference can also inhibit the diffusion of solute atoms into the solvent, which is beneficial to the formation of nanocrystals. Therefore, the atomic size difference can be used as an alternative for the key thermodynamic features affecting alloy properties.

[0054] By determining the importance parameters corresponding to multiple thermodynamic features respectively based on the machine learning algorithm as described above, further analysis can determine that the atomic size difference is the thermodynamic feature with the largest importance parameter.

[0055] Step S408: For each sample, determine the eigenvalue of the sample under the target thermodynamic feature, and add the eigenvalue to the sample to obtain a second intermediate data sample set.

[0056] Calculate the eigenvalue of the sample under the target thermodynamic feature according to the following formula:

[0057]

[0058] where δr represents the eigenvalue of the sample under the atomic size difference feature; c i and c j are the atomic percentages corresponding to the i-th element and the j-th element in the sample respectively, r i represents the atomic radius of the i-th element; n represents the number of elements in the sample; represents the average atomic radius of the n elements in the sample.

[0059] After determining the key features that have the greatest impact on alloy properties, the features are standardized. Feature standardization is to uniformly scale the features to a specified range to eliminate the influence of features with different orders of magnitude on the machine learning model. Therefore, in this embodiment, the Z-score standardization method is adopted, where the ranges of the mean and variance are 0 and 1 respectively. The steps of performing feature standardization processing on the samples in the second intermediate data sample set are as follows: Perform feature standardization processing on the samples in the second intermediate data sample set according to the following formula:

[0060] i = 1, 2, …, M; j = 1, 2, …, N;

[0061] where, N represents the number of samples; M represents the number of features in the samples; represents the feature after standardizing the i-th feature in the j-th sample; represents the initial feature of the i-th feature in the j-th sample; μ (i) represents the average value corresponding to the i-th feature among N samples; σ (i) represents the standard deviation corresponding to the i-th feature among N samples.

[0062] It should be noted that the atomic size difference discussed in the embodiments of the present application can also be replaced with other relevant material knowledge. This dataset construction method can be used not only in the design of aluminum alloys, but also in the design of other materials, including copper alloys, titanium alloys, high-entropy alloys, steels, semiconductor materials, batteries, and so on.

[0063] The embodiments of the present application also provide a method for constructing an alloy data sample set, which is a new dataset construction method applied to machine learning modeling. It proposes to add key thermodynamic features of materials on the basis of the original dataset, rather than just modeling the composition and heat treatment system, realizing the screening of key features in the dataset, which is beneficial to the modeling of machine learning. In this embodiment, key features are added based on factors such as the thermodynamic mechanism and strengthening mechanism of the material itself, improving the interpretability of the machine learning model, and thus enhancing the accuracy of the prediction of the machine learning model.

[0064] Based on the above method embodiments, the embodiments of the present application also provide a device for constructing an alloy data sample set, as shown in Figure 5 shown, the device includes:

[0065] An initial sample set acquisition module 52, configured to acquire an initial alloy data sample set; the samples in the initial alloy data sample set include alloy preparation strategies and the alloy properties corresponding to the alloy preparation strategies; wherein, the alloy preparation strategies include: alloy element content and heat treatment parameters; a data cleaning module 54, configured to clean the samples in the initial alloy data sample set to obtain a first intermediate data sample set; a feature screening and adding module 56, configured to screen and add features to the samples in the first intermediate data sample set to obtain a second intermediate data sample set; a feature standardization module 58, configured to perform feature standardization processing on the samples in the second intermediate data sample set to obtain a target alloy data sample set corresponding to the initial alloy data sample set.

[0066] The above initial sample set acquisition module 52 is configured to screen target documents describing alloy casting processes from the literature on aluminum alloys; screen data including the corresponding relationship between alloy preparation strategies and alloy properties from the target documents; and establish an initial alloy data sample set based on the data including the corresponding relationship between alloy preparation strategies and alloy properties.

[0067] In a preferred embodiment of the present application, the above data cleaning module 54 is configured to extract alloy elements that appear a number of times reaching a number threshold and have a content reaching a content threshold from the initial alloy data sample set, extract heat treatment parameters with both solution treatment and aging, and extract alloy properties including at least one of tensile strength, yield strength, and elongation; and generate a first intermediate data sample set based on the extracted alloy elements and corresponding contents, heat treatment parameters, and alloy properties.

[0068] In a preferred embodiment of the present application, the above feature screening and adding module 56 is configured to obtain a plurality of thermodynamic features that affect the mechanical properties of the alloy; the plurality of thermodynamic features include: mixing enthalpy, mixing entropy, electronegativity, lattice distortion energy, valence electron concentration; determine importance parameters corresponding to the plurality of thermodynamic features respectively based on a machine learning algorithm; determine the thermodynamic feature corresponding to the largest importance parameter as the target thermodynamic feature; for each sample, determine the feature value of the sample under the target thermodynamic feature, and add the feature value to the sample to obtain a second intermediate data sample set.

[0069] In a preferred embodiment of the present application, the above plurality of thermodynamic features include: mixing enthalpy, mixing entropy, electronegativity, lattice distortion energy, valence electron concentration, atomic size difference; the target thermodynamic feature is atomic size difference.

[0070] In a preferred embodiment of the present application, the above feature screening and adding module 56 is configured to calculate the feature value of the sample under the target thermodynamic feature according to the following formula:

[0071]

[0072] wherein, δr represents the feature value of the sample under the atomic size difference feature; c i and c j are the atomic percentages corresponding to the i-th element and the j-th element in the sample respectively, r i represents the atomic radius of the i-th element; n represents the number of elements in the sample; represents the average atomic radius of the n elements in the sample.

[0073] In a preferred embodiment of the present application, the above feature standardization module 58 is configured to perform feature standardization processing on the samples in the second intermediate data sample set according to the following formula:

[0074] i = 1, 2,..., M; j = 1, 2,..., N;

[0075] wherein, N represents the number of samples; M represents the number of features in the sample; denotes the feature after standardization processing for the \(i\)-th feature in the \(j\)-th sample; denotes the initial feature of the \(i\)-th feature in the \(j\)-th sample; \(\mu\) (i) denotes the average value corresponding to the \(i\)-th feature among \(N\) samples; \(\sigma\) (i) denotes the standard deviation corresponding to the \(i\)-th feature among \(N\) samples.

[0076] The device provided by the embodiment of the present application has the same implementation principle and the same technical effects as those of the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the embodiment of the device, reference may be made to the corresponding content in the foregoing method embodiment.

[0077] The embodiment of the present application also provides an electronic device, as Figure 6 shown, which is a schematic structural diagram of the electronic device. Among them, the electronic device includes a processor 61 and a memory 60. The memory 60 stores computer-executable instructions that can be executed by the processor 61, and the processor 61 executes the computer-executable instructions to implement the above method.

[0078] In Figure 6 the shown embodiment, the electronic device further includes a bus 62 and a communication interface 63. Among them, the processor 61, the communication interface 63, and the memory 60 are connected through the bus 62.

[0079] Among them, the memory 60 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 63 (which can be wired or wireless), a communication connection is realized between the system network element and at least one other network element. The Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 62 can be an ISA (Industry Standard Architecture, industrial standard architecture) bus, a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus, or an EISA (Extended Industry Standard Architecture, extended industrial standard architecture) bus, etc. The bus 62 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 only a bidirectional arrow is used in

[0080] The processor 61 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 61 or the instructions in the form of software. The above-mentioned processor 61 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor 61 reads the information in the memory and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0081] The embodiments of the present application also provide a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions cause the processor to implement the above method. For the specific implementation, reference may be made to the foregoing method embodiments and will not be elaborated herein.

[0082] The computer program product of the method, apparatus and electronic device provided by the embodiments of the present application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For the specific implementation, reference may be made to the method embodiments and will not be elaborated herein.

[0083] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present application.

[0084] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0085] In the description of this application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this application. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0086] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of this application, used to illustrate the technical solutions of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in this application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for constructing an alloy data sample set, characterized in that, the method comprises: Obtaining an initial alloy data sample set; the samples in the initial alloy data sample set include alloy preparation strategies and the alloy properties corresponding to the alloy preparation strategies; wherein, the alloy preparation strategies include: alloy element contents and heat treatment parameters; Performing data cleaning on the samples in the initial alloy data sample set to obtain a first intermediate data sample set, including: extracting alloy elements that appear a number of times reaching a number threshold and have a content reaching a content threshold from the initial alloy data sample set, extracting heat treatment parameters with both solution treatment and aging, and extracting alloy properties including at least one of tensile strength, yield strength, and elongation; generating a first intermediate data sample set based on the extracted alloy elements and corresponding contents, heat treatment parameters, and alloy properties; Performing feature screening and feature addition on the samples in the first intermediate data sample set to obtain a second intermediate data sample set, including: obtaining a plurality of thermodynamic features affecting the mechanical properties of the alloy; the plurality of thermodynamic features include: mixing enthalpy, mixing entropy, electronegativity, lattice distortion energy, valence electron concentration; determining importance parameters respectively corresponding to the plurality of thermodynamic features based on a machine learning algorithm; determining the thermodynamic feature corresponding to the largest importance parameter as the target thermodynamic feature; the target thermodynamic feature is the atomic size difference feature; for each sample, determining the feature value of the sample under the target thermodynamic feature and adding the feature value to the sample to obtain the second intermediate data sample set; wherein, the step of determining the feature value of the sample under the target thermodynamic feature includes: calculating the feature value of the sample under the target thermodynamic feature according to the following formula: ; ; Among them, represents the eigenvalue of the sample under the atomic size difference feature; is the atomic percentage corresponding to the th element in the sample, represents the th atomic radius of the element; represents the number of elements in the sample; represents the average atomic radius of the th elements in the sample; Performing feature standardization processing on the samples in the second intermediate data sample set to obtain the target alloy data sample set corresponding to the initial alloy data sample set, including: performing feature standardization processing on the samples in the second intermediate data sample set according to the following formula: ; ; Among them, ; ; represents the number of samples; represents the number of features in the sample; represents the th feature in the th sample after standardization; represents the th feature in the th sample, the initial feature; represents the average value corresponding to the th feature in the represents the standard deviation corresponding to the th feature in the 2. The method according to claim 1, characterized in that, the step of obtaining the initial alloy data sample set includes: Screening target documents describing alloy casting processes from documents on aluminum alloys; Screening data containing the corresponding relationship between alloy preparation strategies and alloy properties from the target documents; Based on the data containing the corresponding relationship between alloy preparation strategies and alloy properties, establishing an initial alloy data sample set.

3. An apparatus for constructing an alloy data sample set, characterized in that, the apparatus comprises: An initial sample set obtaining module, configured to obtain an initial alloy data sample set; the samples in the initial alloy data sample set include alloy preparation strategies and the alloy properties corresponding to the alloy preparation strategies; wherein, the alloy preparation strategies include: alloy element contents and heat treatment parameters; A data cleaning module, which is used to clean the samples in the initial alloy data sample set to obtain a first intermediate data sample set, including: extracting alloy elements that appear a number of times reaching a number threshold and have a content reaching a content threshold from the initial alloy data sample set, extracting heat treatment parameters with both solution treatment and aging, and extracting alloy properties including at least one of tensile strength, yield strength, and elongation; generating a first intermediate data sample set based on the extracted alloy elements and corresponding contents, heat treatment parameters, and alloy properties; A feature screening and adding module, which is used to screen and add features to the samples in the first intermediate data sample set to obtain a second intermediate data sample set, including: obtaining a plurality of thermodynamic features that affect the mechanical properties of the alloy; the plurality of thermodynamic features include: mixing enthalpy, mixing entropy, electronegativity, lattice distortion energy, valence electron concentration; determining importance parameters corresponding to the plurality of thermodynamic features respectively based on a machine learning algorithm; determining the thermodynamic feature corresponding to the largest importance parameter as the target thermodynamic feature; the target thermodynamic feature is the atomic size difference feature; for each sample, determining the feature value of the sample under the target thermodynamic feature, and adding the feature value to the sample to obtain the second intermediate data sample set; wherein, the step of determining the feature value of the sample under the target thermodynamic feature includes: calculating the feature value of the sample under the target thermodynamic feature according to the following formula: ; ; Among them, represents the eigenvalue of the sample under the atomic size difference feature; is the atomic percentage corresponding to the th element in the sample, represents the th atomic radius of the element; represents the number of elements in the sample; represents the th average atomic radius of the elements in the sample; A feature standardization module, which is used to perform feature standardization processing on the samples in the second intermediate data sample set to obtain a target alloy data sample set corresponding to the initial alloy data sample set, including: performing feature standardization processing on the samples in the second intermediate data sample set according to the following formula: ; ; Among them, ; ; represents the number of samples; represents the number of features in the sample; represents the th feature in the th sample after standardization; represents the th feature in the th sample, which is the initial feature; represents the th sample, the average value corresponding to the th feature; represents the th sample, the standard deviation corresponding to the th feature.

4. An electronic device, characterized in that, it includes a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method according to any one of claims 1 to 2.

5. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • Construction method of high-temperature alloy grain size identification model and size identification method

    CN112836433A

  • FULL-VIEW-FIELD QUANTITATIVE STATISTICAL DISTRIBUTION REPRESENTATION METHOD FOR MICROSTRUCTURES of y' PHASES IN METAL MATERIAL

    US20210033549A1