A data processing method for R&D management of new engineering materials
By obtaining the importance index and correlation coefficient of the dimension, performing random classification and classification evaluation, screening the optimal classification results, and adjusting the dimension sequence, the problem that the principal component analysis algorithm does not consider the actual R&D focus direction and dimensional correlation, and achieving more effective dimensionality reduction analysis.
Patent Information
- Application Number
- CN202510607773.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing principal component analysis algorithm does not consider the correlation between the actual R&D direction and dimension in the process of reducing the dimensionality of engineering new materials R&D data, resulting in poor analysis results.
By obtaining the importance index and correlation coefficient of each dimension, perform random classification and classification evaluation, filter the optimal classification results, and adjust the dimension sequence based on the optimal classification results to perform dimension reduction.
The effect of dimensionality reduction analysis of new engineering materials research and development data has been improved, ensuring that the dimension reflects the focus direction of R&D and has less information loss after dimensionality reduction.
Smart Images

Figure CN120145026B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of new material research and development data analysis, and in particular to a data processing method for the management of new engineering material research and development. Background Art
[0002] During the research and development of new engineering materials, a large amount of research and development data of different types or dimensions is usually generated. The research and development data between many different dimensions are also correlated, which increases the complexity of problem analysis. If the research and development data of each dimension are analyzed separately, the analysis results are often isolated and the information in the research and development data cannot be fully utilized, resulting in erroneous analysis results.
[0003] In related technologies, principal component analysis algorithms are usually used to perform dimensionality reduction analysis on the dimensions of engineering new material R&D data. While reducing the number of analysis dimensions and reducing the complexity of analysis, it can also ensure that the information loss contained in the original R&D data is small, so as to achieve the purpose of comprehensive analysis of the R&D process of new materials. However, since the existing principal component analysis algorithm only performs dimensionality reduction analysis based on statistical characteristics such as the variance contribution rate of R&D data, it does not take into account the specific R&D focus of new materials in the actual R&D process, as well as the correlation between the R&D data of various dimensions of new materials, resulting in poor effect of dimensionality reduction analysis of new material R&D data. Summary of the Invention
[0004] In order to solve the technical problem that the existing principal component analysis algorithm only performs dimensionality reduction analysis based on statistical characteristics such as the variance contribution rate of R&D data, and does not take into account the specific R&D focus of new engineering materials in the actual R&D process, as well as the correlation between the R&D data of various dimensions of new materials, resulting in poor analysis of the R&D data of new engineering materials, the purpose of the present invention is to provide a data processing method for the R&D management of new engineering materials. The technical solutions adopted are as follows:
[0005] The present invention proposes a data processing method for the research and development management of new engineering materials, the method comprising:
[0006] Obtain R&D data of different dimensions in the process of new material R&D, and obtain the importance index of each dimension based on the R&D data of different dimensions;
[0007] Based on the R&D data of any two dimensions, a correlation coefficient between the two dimensions is obtained; taking any one dimension as a target dimension, the priority of the target dimension is obtained according to the difference in importance index between the target dimension and each other dimension, the correlation coefficient between the target dimension and each other dimension, and the importance index of the target dimension;
[0008] All dimensions are randomly classified to obtain different classification results, each classification result includes at least two clusters, and each cluster includes at least two dimensions. Any classification result is used as the target classification result. According to the difference in the priority of each dimension in different clusters in the target classification result, the difference in the priority of each dimension in the same cluster, and the correlation coefficient, a classification evaluation parameter of the target classification result is obtained. Based on the classification evaluation parameter, the optimal classification result is screened out from all classification results, and the cluster in the optimal classification result is used as the optimal cluster.
[0009] All dimensions are sorted according to their priority to obtain an initial dimension sequence; based on the optimal cluster of each dimension, the position of the dimension in the initial dimension sequence is adjusted to obtain different priority dimension sequences; based on the priority of the dimension in each priority dimension sequence, the research and development data of the dimension, and the correlation coefficient between the dimensions, a sequence evaluation parameter of each priority dimension sequence is obtained, and the priority dimension sequence corresponding to the maximum value of the sequence evaluation parameter is used as the optimal dimensionality reduction sequence;
[0010] The research and development data of new materials are reduced in dimension based on the optimal dimensionality reduction sequence.
[0011] Furthermore, obtaining the priority of the target dimension according to the difference in importance index between the target dimension and each other dimension, the correlation coefficient between the target dimension and each other dimension, and the importance index of the target dimension includes:
[0012] Negatively correlate the absolute value of the difference in importance index between the target dimension and each other dimension to obtain the adjustment parameter between the target dimension and each other dimension;
[0013] taking the product value of the adjustment parameter and the correlation coefficient between the target dimension and each other dimension as the adjusted correlation coefficient between the target dimension and each other dimension;
[0014] The independence parameter of the target dimension was obtained by normalizing the average of the adjusted correlation coefficients between the target dimension and all other dimensions to negative correlations;
[0015] The product value of the independence parameter of the target dimension and the importance index is used as the priority of the target dimension.
[0016] Furthermore, the random classification of all dimensions to obtain different classification results includes:
[0017] Set a preset number of sets and initialize each set to an empty set;
[0018] Traverse all dimensions and randomly assign each dimension to any set. After the traversal is completed, the non-empty set is regarded as a cluster, and the combination of all clusters is regarded as a classification result.
[0019] All dimensions are randomly classified in the same way to obtain different classification results. When no new classification results can be obtained, the random classification is stopped to obtain different classification results.
[0020] Furthermore, the classification evaluation parameters of the target classification results are obtained based on the difference in priority of each dimension in different clusters, the difference in priority of each dimension in the same cluster, and the correlation coefficient, including:
[0021] In the target classification results, the average priority of all dimensions in each cluster is taken as the overall priority of each cluster;
[0022] performing negative correlation normalization on the variance of the overall priority of all clusters in the target classification result to obtain a first classification evaluation index of the target classification result;
[0023] Taking any cluster in the target classification result as the target cluster, negatively correlating each dimension in the target cluster with the average value of the correlation coefficient between each dimension in other clusters to obtain the inter-cluster independence of the target cluster;
[0024] The combination of any two dimensions in the target cluster is used as the dimension group to be tested, and the absolute value of the difference in priority between the two dimensions in the dimension group to be tested is used as the priority difference between the two dimensions in the dimension group to be tested;
[0025] Multiplying the priority difference between two dimensions in the dimension group to be measured by the correlation coefficient between the two dimensions to obtain a modified correlation coefficient of the dimension group to be measured;
[0026] The average value of the corrected correlation coefficients of all the dimension groups to be tested in the target cluster is taken as the intra-cluster correlation of the target cluster;
[0027] The product value of the inter-cluster independence and the intra-cluster correlation is used as the correlation parameter of the target cluster; the average value of the correlation parameters of all clusters in the target classification result is used as the second classification evaluation index of the target classification result;
[0028] The product value of the first classification evaluation index and the second classification evaluation index is used as the classification evaluation parameter of the target classification result.
[0029] Furthermore, the selecting the best classification result from all classification results based on the classification evaluation parameters includes:
[0030] The classification result corresponding to the maximum value of the classification evaluation parameter is taken as the optimal classification result.
[0031] Furthermore, the position of each dimension in the initial dimension sequence is adjusted based on the optimal cluster of each dimension to obtain different priority dimension sequences, including:
[0032] The number of optimal clusters is used as the reference number;
[0033] In the initial dimension sequence, the sequence composed of the first reference number of dimensions is used as the main sequence, and the sequence composed of the other dimensions except the first reference number of dimensions is used as the secondary sequence;
[0034] If the optimal clusters of the dimensions in the primary sequence are the same, the dimensions in the secondary sequence are traversed, and the first dimension that has not been exchanged and is different from the optimal cluster of each dimension in the primary sequence is selected as the secondary exchange dimension; among all the dimensions in the primary sequence with the same optimal cluster, the dimension closest to the secondary exchange dimension is selected as the primary exchange dimension; the positions of the primary exchange dimension and the secondary exchange dimension are exchanged to obtain the adjusted dimension sequence after each exchange;
[0035] If there is no dimension in the primary sequence that has the same optimal cluster, then traverse the dimensions in the secondary sequence and select the first dimension that has not been exchanged as the secondary exchange dimension; select the dimension in the primary sequence that has the same optimal cluster as the secondary exchange dimension as the primary exchange dimension; swap the positions of the primary exchange dimension and the secondary exchange dimension to obtain an adjusted dimension sequence after each exchange;
[0036] The adjusted dimension sequence after each exchange is divided into a main sequence and a secondary sequence in the same way. If there are still dimensions that have not been exchanged in the secondary sequence, the adjusted dimension sequence is used as the new initial dimension sequence and the dimensions are continued to be exchanged. Otherwise, the exchange is stopped.
[0037] The initial dimension sequence and the adjusted dimension sequence obtained each time are used as the priority dimension sequence.
[0038] Furthermore, obtaining the sequence evaluation parameters of each priority dimension sequence according to the priority of the dimension in each priority dimension sequence, the R&D data of the dimension, and the correlation coefficient between the dimensions includes:
[0039] Normalizing the accumulated values of the priority levels of all dimensions in the main sequence of each priority dimension sequence to obtain a first sequence evaluation index for each priority dimension sequence;
[0040] In the main sequence of each priority dimension sequence, a combination of any two dimensions is used as a dimension group to be analyzed, and the average value of the correlation coefficient between the two dimensions in all dimension groups to be analyzed is negatively correlated to obtain a second sequence evaluation index of each priority dimension sequence;
[0041] Based on the principal component analysis algorithm, the variance contribution rate of each dimension is obtained according to the R&D data of each dimension;
[0042] The cumulative value of the variance contribution rate of all dimensions in the main sequence of each priority dimension sequence is used as the third sequence evaluation index of each priority dimension sequence;
[0043] The sum of the first sequence evaluation index, the second sequence evaluation index and the third sequence evaluation index is used as a sequence evaluation parameter of each priority dimension sequence.
[0044] Furthermore, the dimensionality reduction of the research and development data of new materials based on the optimal dimensionality reduction sequence includes:
[0045] Dimensions are selected one by one from the starting position of the optimal dimensionality reduction sequence, and the cumulative value of the variance contribution rate of all the selected dimensions is used as the numerator, the cumulative value of the variance contribution rate of all the dimensions is used as the denominator, and the ratio is used as the judgment parameter until the judgment parameter is greater than a preset threshold, and the selection is stopped;
[0046] Keep the R&D data of the selected dimension and delete the R&D data of other dimensions.
[0047] Furthermore, the importance index of each dimension obtained based on the R&D data of different dimensions includes:
[0048] Obtaining R&D data on various dimensions of materials other than the new material through big data technology, and obtaining the R&D focus of each other material, scoring each dimension of the other materials in combination with the R&D focus to obtain a score for each dimension of the other materials, and inputting the R&D data and scores of each dimension of the other materials into a neural network for training to obtain a trained neural network;
[0049] Obtain the research and development focus of the new material, input the research and development data and research and development focus of all dimensions of the new material into the trained neural network, and obtain the importance index of each dimension.
[0050] Furthermore, obtaining the correlation coefficient between any two dimensions based on the R&D data of any two dimensions includes:
[0051] The absolute value of the Pearson correlation coefficient between the R&D data of any two dimensions is taken as the correlation coefficient between any two dimensions.
[0052] The present invention has the following beneficial effects:
[0053] Since the existing principal component analysis algorithm only realizes dimensionality reduction analysis through the statistical characteristics of the data, and does not take into account the focus of the research and development process of new materials, as well as the correlation between the research and development data of each dimension in the research and development process, the present invention first analyzes the acquired research and development data of different dimensions, and reflects the importance of each dimension through the acquired importance index, while also reflecting the key research and development direction of the new material. Considering that in the process of dimensionality reduction of the new material research and development data, it is necessary not only to ensure that each dimension after dimensionality reduction can reflect the research and development direction of the new material, but also to ensure that the information loss contained in the research and development data of each dimension after dimensionality reduction is small, that is, the correlation between the research and development data of each dimension after dimensionality reduction is weak. Therefore, the obtained priority can reflect the possibility of each dimension being retained in the subsequent dimensionality reduction process. Taking into account the need for a certain degree of independence between the R&D data of dimensions with similar priorities, the present invention randomly classifies each dimension and evaluates each classification result based on the obtained classification evaluation parameters, thereby screening out the optimal classification result, which is convenient for improving the effect of dimensionality reduction of R&D data, and then based on the optimal cluster of each dimension, adjust the initial dimension sequence, and screen out the optimal dimensionality reduction sequence under the R&D focus direction of new materials based on the sequence evaluation parameters, and reduce the dimensionality of the R&D data based on the optimal dimensionality reduction sequence, thereby improving the effect of dimensionality reduction analysis of new material R&D data. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0055] Figure 1 A flow chart of a data processing method for R&D management of new engineering materials provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0056] To further illustrate the technical means and effectiveness of the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail a data processing method for the research and development management of new engineering materials, including its specific implementation, structure, features, and effectiveness. In the following description, references to different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0057] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0058] The following describes in detail a specific scheme of a data processing method for research and development management of new engineering materials provided by the present invention with reference to the accompanying drawings.
[0059] See also Figure 1 , which shows a flow chart of a data processing method for R&D management of new engineering materials provided by one embodiment of the present invention, the method comprising:
[0060] Step S1: Acquire R&D data of different dimensions in the process of new material R&D, and obtain the importance index of each dimension based on the R&D data of different dimensions.
[0061] In the research and development process of new materials, a large amount of research and development data of different dimensions is usually generated. At the same time, the research and development data of many different dimensions also have a certain correlation. In related technologies, the principal component analysis algorithm is usually used to perform dimensionality reduction analysis on the dimensions of new material research and development data. While reducing the number of analysis dimensions and reducing the complexity of analysis, it can also ensure that the information loss contained in the original research and development data is small, so as to achieve the purpose of comprehensive analysis of the research and development process of new materials. However, since the existing principal component analysis algorithm only performs dimensionality reduction analysis based on statistical characteristics such as the variance contribution rate of the research and development data of each dimension, it does not take into account the specific research and development focus of new materials in the actual research and development process, as well as the correlation between the research and development data of each dimension of new materials, resulting in poor effect of dimensionality reduction analysis of new material research and development data. Therefore, an embodiment of the present invention proposes a data processing method for engineering new material research and development management to solve this problem.
[0062] The embodiment of the present invention first obtains R&D data of different dimensions generated in the process of new material R&D, where the different dimensions may include, for example, the density, magnetism, conductivity, thermal conductivity, chemical composition, element content or cost of the new material, etc. The specific dimensions need to be selected according to the actual R&D test of the new material, and the amount of R&D data contained in each dimension is equal, so as to facilitate the subsequent analysis of the correlation between the dimensions. At the same time, the collected R&D data of each dimension are cleaned to fill in missing values and repair outliers to ensure the quality of the R&D data. Data cleaning technology is a technical means well known to those skilled in the art and will not be elaborated here.
[0063] After collecting R&D data in multiple dimensions, it is also necessary to determine the current R&D focus of the new material so that the importance of the R&D data in each dimension can be evaluated based on the R&D focus. The R&D focus of the new material is determined by the purpose of the new material R&D, for example, the new material focuses on high temperature resistance or non-oxidation.
[0064] Since new materials have a clear research and development focus during the research and development process, the importance of research and development data in different dimensions to the research and development process of new materials varies. Therefore, based on the research and development data in different dimensions, the importance index of each dimension can be obtained, and the importance index can be used to preliminarily reflect the importance of each dimension to the research and development process of new materials.
[0065] Preferably, in one embodiment of the present invention, the method for obtaining the importance index of each dimension specifically includes:
[0066] The research and development data of various dimensions of multiple materials other than the new material are obtained through big data technology, and the research and development focus of each other material is obtained. The dimensions of each other material are manually scored in combination with the research and development focus to obtain the scores of each dimension of the other materials. The score values are: 0.1, 0.2, 0.3, ..., 1, a total of 10 numerical types, corresponding to 10 different levels of importance. In other embodiments, different types of scores can also be set in other ways, but it must be ensured that the different scores set cannot exceed the value 1 for subsequent processing. The result of the scoring is the importance index of the dimension under the research and development focus of the material. The larger the importance index, the more important the dimension. The more important the degree is under the research and development focus of the material, and then the research and development data and scores of each dimension of other materials are input into a 5-layer fully connected neural network. In other embodiments, neural networks of other structures can also be set, which are not limited here. The cross entropy function is selected as the loss function, and the gradient descent method is used for training until the loss function converges, thereby completing the training of the neural network; and the research and development data of all dimensions of the new material and the research and development focus are input into the trained neural network to obtain the importance index of each dimension, where the importance index of each dimension is the set scores, for example, the importance index of one dimension is 0.1, the importance index of another dimension is 0.2, and so on.
[0067] After obtaining the importance index of each dimension in the new material R&D process, we can subsequently conduct a preliminary evaluation and analysis of the importance of the R&D data of each dimension based on the importance index. At the same time, the importance index can also reflect the focus of the R&D of new materials to a certain extent, making it easier to subsequently combine the importance parameters to reduce the dimensionality of the R&D data.
[0068] Step S2: Based on the R&D data of any two dimensions, obtain the correlation coefficient between any two dimensions; take any one dimension as the target dimension, and obtain the priority of the target dimension based on the difference in importance index between the target dimension and each other dimension and the correlation coefficient between the target dimension and each other dimension.
[0069] Considering that in the process of dimensionality reduction of new material R&D data, it is necessary not only to ensure that the dimensions retained after dimensionality reduction can reflect the R&D focus of new materials, but also to ensure that the R&D data of each dimension retained after dimensionality reduction contain less information loss, that is, the correlation between the R&D data of each dimension retained after dimensionality reduction is weak, or the independence of the R&D data of the retained dimensions is strong, the embodiment of the present invention first obtains the correlation coefficient between any two dimensions based on the R&D data of any two dimensions, and reflects the correlation between the two dimensions through the correlation coefficient, which facilitates the subsequent accurate analysis of the independence of each dimension based on the correlation coefficient between the two dimensions.
[0070] Preferably, in one embodiment of the present invention, the absolute value of the Pearson correlation coefficient between the R&D data of any two dimensions is used as the correlation coefficient between any two dimensions. The value range of the correlation coefficient is The closer the correlation coefficient is to 1, the more correlated the two dimensions are, and the closer it is to 0, the less correlated the two dimensions are, that is, the stronger the independence between the two dimensions is. The calculation method of the Pearson correlation coefficient is a technical means well known to those skilled in the art and will not be elaborated here.
[0071] After obtaining the correlation coefficient between any two dimensions, in order to facilitate analysis, any dimension can be used as the target dimension. If the correlation coefficient between the target dimension and other dimensions is higher, it indicates that the R&D data of the target dimension can be reflected by the R&D data of other dimensions, and the target dimension is less likely to be retained in the process of dimensionality reduction analysis. At the same time, if the importance index of the target dimension is smaller, it indicates that the connection between the target dimension and the R&D focus of the new material is weaker, and the target dimension is less likely to be retained in the process of dimensionality reduction analysis. At the same time, in order to avoid mutual interference between dimensions with large differences in importance index, it is also necessary to adjust the correlation between dimensions based on the difference in importance index between the target dimension and other dimensions. Therefore, the difference in importance index between the target dimension and each other dimension, the correlation coefficient between the target dimension and each other dimension, and the importance index of the target dimension can be analyzed. The priority obtained reflects the possibility of each dimension being retained in the dimensionality reduction process, which facilitates a more effective dimensionality reduction analysis of the R&D data of new materials based on the priority.
[0072] Preferably, in one embodiment of the present invention, the method for obtaining the priority of the target dimension specifically includes:
[0073] The absolute value of the difference in importance index between the target dimension and each other dimension is negatively correlated to obtain the adjustment parameter between the target dimension and each other dimension; the product of the adjustment parameter and the correlation coefficient between the target dimension and each other dimension is used as the adjusted correlation coefficient between the target dimension and each other dimension; the average value of the adjusted correlation coefficient between the target dimension and all other dimensions is negatively normalized to obtain the independence parameter of the target dimension; the product of the independence parameter of the target dimension and the importance index is used as the priority of the target dimension. The expression of priority can be specifically, for example:
[0074]
[0075] in, Indicates the priority of the target dimension; Indicates the importance index of the target dimension; Indicates the dimension other than the target dimension. Importance index of other dimensions; Indicates the target dimension and the The correlation coefficients between the other dimensions; represents the number of all dimensions, then Indicates the number of dimensions other than the target dimension; Expressed as a natural constant An exponential function with base .
[0076] In the process of obtaining the priority of the target dimension, the priority The larger the value, the greater the possibility that the target dimension will be retained in the subsequent dimensionality reduction process. At the same time, the priority can also reflect the independence of the target dimension relative to other dimensions and the importance of the target dimension in the new material research and development process. The importance index of the target dimension is The larger the value, the stronger the connection between the target dimension and the research and development focus of the new material, and the more important the target dimension is in the research and development process, the greater the possibility that the target dimension will be retained in the subsequent dimensionality reduction process, that is, the priority At the same time, it is also necessary to ensure that the information loss contained in the R&D data of each dimension retained after dimensionality reduction is small, that is, the correlation between the R&D data of each dimension retained after dimensionality reduction is weaker and the independence is stronger, so the correlation coefficient between the target dimension and other dimensions is The smaller it is, the stronger the independence of the target dimension is, and the greater the possibility that the target dimension will be retained in the subsequent dimensionality reduction process. The larger the R&D focus, the greater the impact on R&D data of different dimensions during the dimensionality reduction process. When calculating the independence of R&D data of each dimension, the calculation ratio of the correlation coefficient between the dimensions should be adjusted to avoid mutual interference between dimensions with large differences in importance index. It reflects the difference in importance index between the target dimension and other dimensions. The larger the difference value, the smaller the proportion of the correlation coefficient between the target dimension and other dimensions in the calculation result should be. Therefore, by adjusting the parameter Correlation coefficient Make adjustments to obtain the adjusted correlation coefficient and use the natural constant The exponential function with the base ∑ is used to normalize the negative correlation between the target dimension and the average of the adjusted correlation coefficients of all other dimensions, and the independence parameter is used. Reflects the degree of independence of the target dimension relative to all other dimensions.
[0077] After obtaining the priority of the target dimension, the priority of each other dimension can be obtained by the same method as above. In the subsequent process, further analysis can be made on each dimension in the new material research and development process based on the priority, thereby achieving dimensionality reduction of the research and development data.
[0078] Step S3: Randomly classify all dimensions to obtain different classification results. Each classification result includes at least two clusters, and each cluster includes at least two dimensions. Any classification result is used as the target classification result. According to the difference in the priority of each dimension in different clusters in the target classification result, the difference in the priority of each dimension in the same cluster, and the correlation coefficient, the classification evaluation parameters of the target classification result are obtained. Based on the classification evaluation parameters, the optimal classification result is screened out from all classification results, and the cluster in the optimal classification result is used as the optimal cluster.
[0079] In the principal component analysis algorithm, the number of principal components, that is, the number of retained dimensions, is usually determined by calculating statistical features such as the variance contribution rate of R&D data of different dimensions. In the R&D process of new materials, different R&D focuses have different emphases on the R&D data of corresponding dimensions. Therefore, the number of dimensions to be retained cannot be directly calculated by statistical features. It is also necessary to analyze in combination with the priority of each dimension. Considering that the R&D data between all retained dimensions must also have a certain degree of relative independence after data dimensionality reduction, the dimensions with higher priority need to be retained in the subsequent process. However, in terms of priority, the relative independence of R&D data between dimensions with similar priority cannot be guaranteed. It is also necessary to analyze the dimensionality reduction process in combination with the relationship between the priorities of each dimension. Therefore, the embodiment of the present invention obtains multiple classification results by randomly classifying all dimensions. Each classification result includes at least two clusters, and each cluster includes at least two dimensions, which facilitates the subsequent screening of the optimal classification result from multiple classification results and facilitates subsequent dimensionality reduction processing.
[0080] Preferably, in one embodiment of the present invention, the method for obtaining different classification results specifically includes:
[0081] Set a preset number of sets, initialize each set to an empty set; traverse all dimensions, and randomly assign each dimension to any set. After the traversal is completed, use the non-empty set as a cluster, and use the combination of all clusters as a classification result. For example: this process can be compared to each set as a basket, and each dimension as a ball of a different color, and then each ball is randomly thrown into one of the baskets; randomly classify all dimensions in the same way to obtain different classification results, until no new classification results can be obtained, then stop random classification, and obtain different classification results. The method to distinguish different classification results is: if the number of clusters in the two classification results is different, then the two classification results are different. If the number of clusters in the two classification results is the same, but the dimensions contained in the corresponding clusters in the two classification results are different, then the two classification results are Different; it should be noted that in the process of random classification of each dimension, it is necessary to ensure that the number of non-empty sets after each random classification is not less than 2, and the number of elements in each non-empty set is not less than 2, so as to subsequently analyze the correlation between dimensions between clusters and the correlation between dimensions within clusters. The preset number is set to the total number of dimensions, and the specific value of the preset number can also be set by the implementer according to the specific implementation scenario, and is not limited here. It should be noted that when the total number of dimensions is an even number, the existence of only two dimensions in each cluster is also a classification result. At this time, the number of clusters is half of the total number of dimensions. Therefore, the number of sets set, that is, the preset number should not be less than half of the total number of dimensions. Combined with the two cases where the total number of dimensions is even and odd, it is necessary to ensure that the preset number is not less than half of the total number of dimensions and rounded down.
[0082] Since it is necessary to combine the priority to determine the dimensions to be retained and the relative independence between the dimensions must also be determined, multiple classification results are obtained by randomly classifying all dimensions. In each classification result, dimensions with different correlation levels are divided into multiple clusters. In the subsequent dimensionality reduction analysis, it is necessary to ensure that the correlation between the R&D data of each dimension within the cluster in the classification result should be large, while the correlation between the R&D data of each dimension between clusters should be small. In addition, in the process of classifying each dimension, the relationship between the R&D focus should also be considered, that is, the relationship between the priority of different dimensions needs to be introduced. Therefore, the embodiment of the present invention takes any classification result as the target classification result, and analyzes the differences in the priority of each dimension in different clusters in the target classification result, the differences in the priority of each dimension in the same cluster, and the correlation coefficients between the dimensions, and evaluates the target classification result through the obtained classification evaluation parameters, so that the optimal classification result can be screened out from all classification results based on the classification evaluation parameters in the subsequent process.
[0083] Preferably, in one embodiment of the present invention, the method for obtaining the classification evaluation parameters of the target classification result specifically includes:
[0084] In the target classification results, the average value of the priority of all dimensions in each cluster is used as the overall priority of each cluster; the variance of the overall priority of all clusters in the target classification results is negatively normalized to obtain the first classification evaluation index of the target classification results; any cluster in the target classification results is used as the target cluster, and the average value of the correlation coefficient between each dimension in the target cluster and each dimension in other clusters is negatively correlated to obtain the inter-cluster independence of the target cluster; the combination of any two dimensions in the target cluster is used as the dimension group to be tested, and the absolute value of the difference in priority between the two dimensions in the dimension group to be tested is used as the first classification evaluation index of the target classification results; The priority difference between two dimensions in the dimension group to be tested; multiplying the priority difference between two dimensions in the dimension group to be tested by the correlation coefficient between the two dimensions to obtain the corrected correlation coefficient of the dimension group to be tested; taking the average value of the corrected correlation coefficients of all dimension groups to be tested in the target cluster as the intra-cluster correlation of the target cluster; taking the product of inter-cluster independence and intra-cluster correlation as the correlation parameter of the target cluster; taking the average value of the correlation parameters of all clusters in the target classification result as the second classification evaluation index of the target classification result; taking the product of the first classification evaluation index and the second classification evaluation index as the classification evaluation parameter of the target classification result. The expression of the classification evaluation parameter can be specifically, for example:
[0085]
[0086]
[0087]
[0088]
[0089] in, Classification evaluation parameters representing target classification results; The first classification evaluation index representing the target classification result; The second classification evaluation index representing the target classification result; Expressed as a natural constant An exponential function with base ; represents the function of taking the variance; Indicates a collection symbol; Indicates the target classification result The average priority of all dimensions in the cluster, that is, The overall priority of each cluster; Indicates the number of all clusters in the target classification result; Indicates the target classification result The correlation parameters of the clusters; Represents the correlation parameter of the target cluster in the target classification result; Indicates the number of dimensions in the target cluster in the target classification result, represents the number of dimensions in clusters other than the target cluster in the target classification result, then , is the number of all dimensions; Indicates the first dimensions and the first dimension in other clusters except the target cluster The correlation coefficients between the dimensions; and Indicates the first The priority of two dimensions in the group of dimensions to be measured; Indicates the first The correlation coefficient between two dimensions in the group of dimensions to be measured; Indicates the number of dimension groups to be measured in the target cluster.
[0090] In the process of obtaining the classification evaluation parameters of the target classification results, the classification evaluation parameters Used to evaluate the target classification results, classification evaluation parameters The larger the value, the better the target classification result. In the target classification result, it is necessary to avoid excessive concentration of dimensions with similar priorities. That is, the distribution of the priority of dimensions between clusters should be as uniform as possible, so that the dimensions retained after subsequent dimensionality reduction can reflect the research and development focus of new materials. Therefore, the variance of the overall priority of all clusters in the target classification result is The smaller it is, the more evenly the priority distribution of dimensions among various clusters in the target classification results is. Therefore, the natural constant The exponential function with base Normalize the negative correlation to obtain the first classification evaluation index , The larger the value is, the better the target classification result is. The larger the value, the greater the correlation of the R&D data of each dimension within the cluster should be, and the correlation of the R&D data of each dimension between clusters should be smaller. That is, the R&D data of each dimension between clusters need to have a certain degree of independence. It can reflect the correlation of each dimension between the target cluster and other clusters, and perform negative correlation mapping to obtain the inter-cluster independence of the target cluster. ,and It can reflect the correlation between the two dimensions within the target cluster. At the same time, in order to avoid the priority of each dimension in the cluster being too different, the priority difference is used. right Make adjustments to obtain the corrected correlation coefficient of the dimension group to be measured , combined with the corrected correlation coefficients of all the dimension groups to be tested, through the intra-cluster correlation of the target cluster Reflects the correlation between the dimensions within the target cluster, and combines the inter-cluster independence and intra-cluster correlation to obtain the correlation parameter , correlation parameter The larger the value is, the greater the correlation between the dimensions in the cluster is, and the more independent the dimensions in the cluster are relative to the dimensions in other clusters. Therefore, the second classification evaluation index is obtained by combining the correlation parameters of all clusters in the target classification results. , The larger the value is, the better the target classification result is. The bigger it is.
[0091] After obtaining the classification evaluation parameters of the target classification result, the classification evaluation parameters of other classification results can be obtained by the same method as above, and then the optimal classification result can be screened out from all classification results based on the classification evaluation parameters. Since the larger the classification evaluation parameter, the better the classification result, in one embodiment of the present invention, the classification result corresponding to the maximum value of the classification evaluation parameter is used as the optimal classification result, and the cluster in the optimal classification result is used as the optimal cluster.
[0092] After obtaining the optimal classification result, the dimension sequence obtained in the subsequent step can be adjusted based on the optimal cluster of each dimension, so as to obtain the optimal dimensionality reduction sequence and realize effective dimensionality reduction of new material research and development data in the dimensionality reduction sequence.
[0093] Step S4: Sort all dimensions according to their priority to obtain an initial dimension sequence; based on the optimal cluster of each dimension, adjust the position of the dimension in the initial dimension sequence to obtain different priority dimension sequences; according to the priority of the dimension in each priority dimension sequence, the research and development data of the dimension, and the correlation coefficient between the dimensions, obtain the sequence evaluation parameters of each priority dimension sequence, and take the priority dimension sequence corresponding to the maximum value of the sequence evaluation parameter as the optimal dimensionality reduction sequence.
[0094] The dimensionality reduction process of the traditional principal component analysis algorithm is to sort each dimension in descending order according to the variance contribution rate of the dimension's R&D data, and then select the dimensions to be retained one by one. Therefore, the traditional principal component analysis algorithm only performs dimensionality reduction analysis based on the statistical characteristics of the data. The dimensions retained after dimensionality reduction using the principal component analysis algorithm cannot reflect the research and development focus of new materials. Since the greater the priority, the greater the possibility that the dimension will be retained in the subsequent dimensionality reduction process, the embodiment of the present invention sorts all dimensions according to the size of the priority to obtain an initial dimension sequence. In one embodiment of the present invention, the dimensions are sorted in descending order of priority. In other embodiments, the dimensions can also be sorted in descending order of priority, which is not limited here. The analysis logic of the two is opposite. Considering that the correlation between dimensions in different optimal clusters is weak, that is, the independence is strong, while the dimensions retained after dimensionality reduction need to have a certain degree of independence, the position of the dimension in the initial dimension sequence can be adjusted based on the optimal cluster in which each dimension belongs to, to obtain different priority dimension sequences. Each priority dimension sequence can be evaluated subsequently to screen the optimal dimensionality reduction sequence.
[0095] Preferably, in one embodiment of the present invention, the method for obtaining different priority dimension sequences specifically includes:
[0096] The number of optimal clusters is taken as the reference number; in the initial dimension sequence, the sequence composed of the first reference number of dimensions is taken as the main sequence, and the sequence composed of the dimensions other than the first reference number of dimensions is taken as the secondary sequence; if the optimal clusters of the dimensions in the main sequence are the same, the dimensions in the secondary sequence are traversed, and the first dimension that has not been exchanged and is different from the optimal clusters of the dimensions in the main sequence is selected as the secondary exchange dimension; among all the dimensions of the same optimal cluster in the main sequence, a dimension closest to the secondary exchange dimension is selected as the main exchange dimension; the positions of the main exchange dimension and the secondary exchange dimension are exchanged to obtain the adjusted dimension sequence after each exchange; if the optimal clusters of the dimensions in the main sequence are not the same, that is, the optimal clusters of the dimensions in the main sequence are different, the dimensions in the secondary sequence are traversed, and the first dimension that has not been exchanged and is different from the optimal clusters of the dimensions in the main sequence is selected. The dimensions that have not been exchanged are used as secondary exchange dimensions; the dimension with the same optimal cluster as the secondary exchange dimension is selected in the main sequence as the main exchange dimension; the positions of the main exchange dimension and the secondary exchange dimension are exchanged to obtain the adjusted dimension sequence after each exchange; the adjusted dimension sequence after each exchange is divided into the main sequence and the secondary sequence in the same way. If there are still dimensions that have not been exchanged in the secondary sequence, the adjusted dimension sequence is used as the new initial dimension sequence, and the dimensions continue to be exchanged, otherwise the exchange is stopped; the initial dimension sequence and the adjusted dimension sequence obtained each time are used as the priority dimension sequence. It should be noted that in the process of adjusting the initial adjustment sequence, the optimal clusters of each dimension in the main sequence tend to be different from each other, so that the independence of the dimensions retained after subsequent dimensionality reduction is strong, and the retained dimensions can reflect the research and development focus of new materials.
[0097] For example: Assume that the number of optimal clusters in the optimal classification result is 2, namely cluster A and cluster B, and the total number of dimensions is 5. The optimal clusters of each dimension in the initial dimension sequence are , where the three dimensions at the 1st, 2nd and 4th positions in the initial dimension sequence are in cluster A, and the two dimensions at the 3rd and 5th positions are in cluster B. In the initial dimension sequence, the first two dimensions are the main sequence and the last three dimensions are the secondary sequence. When the first exchange occurs, since there are dimensions in the same optimal cluster in the main sequence, The corresponding dimension is used as the secondary exchange dimension, The corresponding dimension is used as the main exchange dimension, and the first adjustment dimension sequence after the two are exchanged is , and at this time and The corresponding dimensions are considered to have been exchanged, and the two dimensions in the main sequence of the first adjustment dimension sequence are in different optimal clusters. Therefore, the first adjustment dimension sequence will be exchanged in the next The corresponding dimension is used as the secondary exchange dimension, The corresponding dimension is used as the main exchange dimension, and the second adjustment dimension sequence after the two are exchanged is ,at this time 、 、 and The corresponding dimensions are considered to have been exchanged, and the next exchange will be the second adjustment dimension sequence The corresponding dimension is used as the secondary exchange dimension, The corresponding dimension is used as the main exchange dimension, and the third adjustment dimension sequence after the two are exchanged is At this time, all dimensions in the sequence have been exchanged, so the exchange process ends, and the initial dimension sequence and the three adjusted dimension sequences obtained are used as the priority dimension sequence.
[0098] After obtaining multiple priority dimension sequences, since the dimensions retained after dimensionality reduction of the R&D data of new materials need to not only have strong independence, but also be able to reflect the R&D focus of new materials, the priority of the dimensions in each priority dimension sequence, the R&D data of the dimensions, and the correlation coefficients between the dimensions can be analyzed, and the priority dimension sequence can be evaluated through sequence evaluation parameters to facilitate the subsequent screening of the optimal dimensionality reduction sequence from all priority dimension sequences based on the sequence evaluation parameters.
[0099] Preferably, in one embodiment of the present invention, the method for obtaining the sequence evaluation parameter of each priority dimension sequence specifically includes:
[0100] Normalize the cumulative value of the priority of all dimensions in the main sequence of each priority dimension sequence to obtain the first sequence evaluation index of each priority dimension sequence; in the main sequence of each priority dimension sequence, take the combination of any two dimensions as the dimension group to be analyzed, and negatively map the average value of the correlation coefficient between the two dimensions in all the dimension groups to be analyzed to obtain the second sequence evaluation index of each priority dimension sequence; based on the principal component analysis algorithm, according to the research and development data of each dimension, obtain the variance contribution rate of each dimension. The calculation method of the variance contribution rate is a technical means well known to those skilled in the art and will not be described here; take the cumulative value of the variance contribution rate of all dimensions in the main sequence of each priority dimension sequence as the third sequence evaluation index of each priority dimension sequence; take the sum of the first sequence evaluation index, the second sequence evaluation index and the third sequence evaluation index as the sequence evaluation parameter of each priority dimension sequence. The expression of the sequence evaluation parameter can be specifically, for example:
[0101]
[0102] in, Indicates the Sequence evaluation parameters of priority dimension sequences; Indicates the The first priority dimension sequence in the main sequence The priority of each dimension; Indicates the The number of dimensions in the main sequence of a priority dimension sequence, that is, the number of optimal clusters, is the same as the number of dimensions in the main sequence of different priority dimension sequences; Indicates the The priority of each dimension; Indicates the number of all dimensions; Indicates the The first priority dimension sequence in the main sequence The correlation coefficient between two dimensions in the dimension group to be analyzed; Indicates the The number of dimension groups to be analyzed in the main sequence of a priority dimension sequence is the same as the number of dimension groups to be analyzed in the main sequences of different priority dimension sequences; Indicates the The first priority dimension sequence in the main sequence The variance contribution rate of each dimension.
[0103] In the process of obtaining the sequence evaluation parameters of each priority dimension sequence, the sequence evaluation parameters The larger the value is, the more suitable the priority dimension sequence is for subsequent dimensionality reduction analysis. The first sequence evaluation index is The larger the value is, the greater the proportion of the priority of each dimension in the main sequence of the priority dimension sequence is, which means that the dimensions in the front part of the priority dimension sequence can better reflect the research and development focus of new materials. The more ideal the priority dimension sequence is, the higher the sequence evaluation parameter is. The bigger, Used for Normalize and evaluate the second sequence The larger the value is, the greater the independence of the R&D data of each dimension in the main sequence of the priority dimension sequence is, and the more ideal the priority dimension sequence is. The larger the third order evaluation index is, The larger the value is, the more information the R&D data of each dimension in the main sequence of the priority dimension sequence contains. The less information is lost after dimensionality reduction using the priority dimension sequence, the more ideal the priority dimension sequence is. The sequence evaluation parameter The bigger it is.
[0104] The larger the sequence evaluation parameter is, the more ideal the effect of subsequent dimensionality reduction using the corresponding priority dimension sequence is. Therefore, the priority dimension sequence corresponding to the maximum value of the sequence evaluation parameter can be used as the optimal dimensionality reduction sequence. The independence between the relatively forward dimensions in the optimal dimensionality reduction sequence is strong, and it can reflect the research and development focus of new materials. In the subsequent development, the research and development data of new materials can be analyzed for dimensionality reduction based on the optimal dimensionality reduction sequence.
[0105] Step S5: Reduce the dimensionality of the research and development data of new materials based on the optimal dimensionality reduction sequence.
[0106] The independence of the front dimensions in the optimal dimensionality reduction sequence is relatively strong, and each dimension can reflect the research and development focus of new materials. Therefore, the dimensionality reduction of new material research and development data can be directly performed based on the optimal dimensionality reduction sequence, so that the dimensions retained after dimensionality reduction have a certain degree of independence, and can accurately reflect the research and development focus of new materials, thereby improving the effect of dimensionality reduction of new material research and development data.
[0107] Preferably, in one embodiment of the present invention, the method for obtaining the sequence evaluation parameter of each priority dimension sequence specifically includes:
[0108] In order to reduce the information loss of the R&D data after dimensionality reduction, the dimensions to be retained can be selected in the optimal dimensionality reduction sequence based on the variance contribution rate of the dimension. Therefore, the dimensions can be selected one by one from the starting position of the optimal dimensionality reduction sequence, and the cumulative value of the variance contribution rate of all selected dimensions is used as the numerator, the cumulative value of the variance contribution rate of all dimensions is used as the denominator, and the ratio is used as the judgment parameter. The selection process is stopped until the judgment parameter is greater than the preset threshold. It should be noted that the judgment parameter must be calculated once each time a dimension is selected, and it is judged whether the judgment parameter is greater than the preset threshold; the R&D data of the selected dimension is retained, and the R&D data of other dimensions are deleted, thereby realizing the dimensionality reduction analysis of the new material R&D data. The preset threshold is set to 0.7. The specific value of the preset threshold can also be set by the implementer according to the specific implementation scenario, and is not limited here.
[0109] In summary, the embodiment of the present invention first obtains the research and development data of different dimensions in the process of new material research and development, and uses the trained neural network to obtain the importance index of each dimension; based on the research and development data of any two dimensions, obtains the correlation coefficient between any two dimensions; obtains the priority of each dimension according to the difference and correlation coefficient between the importance index of each dimension and other dimensions and the importance index of each dimension; randomly classifies all dimensions to obtain different classification results, and obtains the classification evaluation parameters of each classification result according to the difference in the priority of each dimension in different clusters in each classification result, the difference in the priority of each dimension in the same cluster and the correlation coefficient, and obtains the classification evaluation parameters based on the classification. The class evaluation parameters filter out the optimal classification result from all classification results, and take the cluster in the optimal classification result as the optimal cluster; sort all dimensions in descending order of priority to obtain an initial dimension sequence; based on the optimal cluster of each dimension, adjust the position of the dimension in the initial dimension sequence to obtain different priority dimension sequences; according to the priority of the dimension in each priority dimension sequence, the research and development data of the dimension, and the correlation coefficient between the dimensions, obtain the sequence evaluation parameters of each priority dimension sequence, and take the priority dimension sequence corresponding to the maximum value of the sequence evaluation parameters as the optimal dimensionality reduction sequence, and then reduce the dimensionality of the research and development data of new materials based on the optimal dimensionality reduction sequence.
[0110] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0111] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A data processing method for the research and development management of new engineering materials, characterized in that: The method comprises: Obtain R&D data of different dimensions in the process of new material R&D, and obtain the importance index of each dimension based on the R&D data of different dimensions; Based on the R&D data of any two dimensions, the correlation coefficient between any two dimensions is obtained; taking any dimension as the target dimension, the priority of the target dimension is obtained based on the difference in importance index between the target dimension and each other dimension, the correlation coefficient between the target dimension and each other dimension, and the importance index of the target dimension; All dimensions are randomly classified to obtain different classification results. Each classification result includes at least two clusters, and each cluster includes at least two dimensions. Any classification result is used as the target classification result. Based on the differences in the priority of each dimension in different clusters in the target classification result, the differences in the priority of each dimension in the same cluster, and the correlation coefficient, the classification evaluation parameters of the target classification result are obtained. Based on the classification evaluation parameters, the optimal classification result is screened out from all classification results, and the cluster in the optimal classification result is used as the optimal cluster. All dimensions are sorted by priority to obtain an initial dimension sequence. Based on the optimal cluster of each dimension, the position of the dimension in the initial dimension sequence is adjusted to obtain different priority dimension sequences. Based on the priority of each dimension in the priority dimension sequence, the research and development data of the dimension, and the correlation coefficient between the dimensions, the sequence evaluation parameter of each priority dimension sequence is obtained. The priority dimension sequence corresponding to the maximum value of the sequence evaluation parameter is used as the optimal dimensionality reduction sequence. Reduce the dimensionality of new material R&D data based on the optimal dimensionality reduction sequence; The classification evaluation parameters for obtaining the target classification result include: In the target classification results, the average priority of all dimensions in each cluster is taken as the overall priority of each cluster; The variance of the overall priority of all clusters in the target classification result is negatively normalized to obtain the first classification evaluation index of the target classification result; Take any cluster in the target classification result as the target cluster, and perform negative correlation mapping on the average value of the correlation coefficient between each dimension in the target cluster and each dimension in other clusters to obtain the inter-cluster independence of the target cluster; The combination of any two dimensions in the target cluster is used as the dimension group to be tested, and the absolute value of the difference in priority between the two dimensions in the dimension group to be tested is used as the priority difference between the two dimensions in the dimension group to be tested; Multiply the priority difference between two dimensions in the dimension group to be measured by the correlation coefficient between the two dimensions to obtain the corrected correlation coefficient of the dimension group to be measured; The average value of the corrected correlation coefficients of all the dimension groups to be tested in the target cluster is taken as the intra-cluster correlation of the target cluster; The product of inter-cluster independence and intra-cluster correlation is used as the correlation parameter of the target cluster; the average value of the correlation parameters of all clusters in the target classification result is used as the second classification evaluation index of the target classification result; The product value of the first classification evaluation index and the second classification evaluation index is used as the classification evaluation parameter of the target classification result; The step of obtaining sequence evaluation parameters for each priority dimension sequence includes: Normalizing the accumulated values of the priority levels of all dimensions in the main sequence of each priority dimension sequence to obtain the first sequence evaluation index of each priority dimension sequence; In the main sequence of each priority dimension sequence, any combination of two dimensions is taken as the dimension group to be analyzed, and the average value of the correlation coefficient between the two dimensions in all the dimension groups to be analyzed is negatively correlated to obtain the second sequence evaluation index of each priority dimension sequence; Based on the principal component analysis algorithm, the variance contribution rate of each dimension is obtained according to the R&D data of each dimension; The cumulative value of the variance contribution rate of all dimensions in the main sequence of each priority dimension sequence is used as the third sequence evaluation index of each priority dimension sequence; The sum of the first sequence evaluation index, the second sequence evaluation index and the third sequence evaluation index is used as the sequence evaluation parameter of each priority dimension sequence.
2. A data processing method for research and development management of new engineering materials according to claim 1, characterized in that: Obtaining the priority of the target dimension according to the difference in importance index between the target dimension and each other dimension, the correlation coefficient between the target dimension and each other dimension, and the importance index of the target dimension includes: Negatively correlate the absolute value of the difference in importance index between the target dimension and each other dimension to obtain the adjustment parameter between the target dimension and each other dimension; taking the product value of the adjustment parameter and the correlation coefficient between the target dimension and each other dimension as the adjusted correlation coefficient between the target dimension and each other dimension; The independence parameter of the target dimension was obtained by normalizing the average of the adjusted correlation coefficients between the target dimension and all other dimensions to negative correlations; The product value of the independence parameter of the target dimension and the importance index is used as the priority of the target dimension.
3. A data processing method for research and development management of new engineering materials according to claim 1, characterized in that: The random classification of all dimensions can obtain different classification results including: Set a preset number of sets and initialize each set to an empty set; Traverse all dimensions and randomly assign each dimension to any set. After the traversal is completed, the non-empty set is regarded as a cluster, and the combination of all clusters is regarded as a classification result. All dimensions are randomly classified in the same way to obtain different classification results. When no new classification results can be obtained, the random classification is stopped to obtain different classification results.
4. A data processing method for research and development management of new engineering materials according to claim 1, characterized in that: The step of selecting the optimal classification result from all classification results based on the classification evaluation parameters includes: The classification result corresponding to the maximum value of the classification evaluation parameter is taken as the optimal classification result.
5. The data processing method for R&D management of new engineering materials according to claim 1, characterized in that: The method of adjusting the position of each dimension in the initial dimension sequence based on the optimal cluster of each dimension to obtain different priority dimension sequences includes: The number of optimal clusters is used as the reference number; In the initial dimension sequence, the sequence composed of the first reference number of dimensions is used as the main sequence, and the sequence composed of the other dimensions except the first reference number of dimensions is used as the secondary sequence; If the optimal clusters of the dimensions in the primary sequence are the same, the dimensions in the secondary sequence are traversed, and the first dimension that has not been exchanged and is different from the optimal cluster of each dimension in the primary sequence is selected as the secondary exchange dimension; among all the dimensions in the primary sequence with the same optimal cluster, the dimension closest to the secondary exchange dimension is selected as the primary exchange dimension; the positions of the primary exchange dimension and the secondary exchange dimension are exchanged to obtain the adjusted dimension sequence after each exchange; If there is no dimension in the primary sequence that has the same optimal cluster, then traverse the dimensions in the secondary sequence and select the first dimension that has not been exchanged as the secondary exchange dimension; select the dimension in the primary sequence that has the same optimal cluster as the secondary exchange dimension as the primary exchange dimension; swap the positions of the primary exchange dimension and the secondary exchange dimension to obtain an adjusted dimension sequence after each exchange; The adjusted dimension sequence after each exchange is divided into a main sequence and a secondary sequence in the same way. If there are still dimensions that have not been exchanged in the secondary sequence, the adjusted dimension sequence is used as the new initial dimension sequence and the dimensions are continued to be exchanged. Otherwise, the exchange is stopped. The initial dimension sequence and the adjusted dimension sequence obtained each time are used as the priority dimension sequence.
6. A data processing method for research and development management of new engineering materials according to claim 5, characterized in that: The dimensionality reduction of the research and development data of new materials based on the optimal dimensionality reduction sequence includes: Dimensions are selected one by one from the starting position of the optimal dimensionality reduction sequence, and the cumulative value of the variance contribution rate of all the selected dimensions is used as the numerator, the cumulative value of the variance contribution rate of all the dimensions is used as the denominator, and the ratio is used as the judgment parameter until the judgment parameter is greater than a preset threshold, and the selection is stopped; Keep the R&D data of the selected dimension and delete the R&D data of other dimensions.
7. A data processing method for research and development management of new engineering materials according to claim 1, characterized in that: The importance index of each dimension obtained based on the R&D data of different dimensions includes: Obtaining R&D data on various dimensions of materials other than the new material through big data technology, and obtaining the R&D focus of each other material, scoring each dimension of the other materials in combination with the R&D focus to obtain a score for each dimension of the other materials, and inputting the R&D data and scores of each dimension of the other materials into a neural network for training to obtain a trained neural network; Obtain the research and development focus of the new material, input the research and development data and research and development focus of all dimensions of the new material into the trained neural network, and obtain the importance index of each dimension.
8. The data processing method for research and development management of new engineering materials according to claim 1, characterized in that: The R&D data based on any two dimensions to obtain the correlation coefficient between any two dimensions includes: The absolute value of the Pearson correlation coefficient between the R&D data of any two dimensions is taken as the correlation coefficient between any two dimensions.
Citation Information
Patent Citations
Resourceful treatment method and system for kitchen waste and municipal sludge
CN118260580A
Material processing cost data analysis method and system for engineering management
CN119831174A