Method for extracting, mining and applying production class data

By segmenting, dimensionality reduction and clustering of electronic manufacturing data, and combining with random forest algorithms to identify key influencing factors, the problem of difficulty in effectively extracting key information in manufacturing data in the existing technology is solved, and the effect of improving production efficiency and reducing production costs is achieved.

CN120179706APending Publication Date: 2025-06-20SHENZHEN KAIFA TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311764518.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Electronic manufacturing data is characterized by large capacity, multiple types, high dimensions, high pass rate and rapid updates. It is difficult for the existing technology to effectively extract the key influencing factors and useful information.

Method used

A method for extraction and mining of production-type data is proposed, including data collection, segmentation according to the data generation time and integration according to specific index names, data characteristics dimensionality reduction and clustering, and random forest algorithm importance sorting to identify key influencing factors and guide production processes.

Benefits of technology

By deeply digging production data, we use slicing, dimensionality reduction and extraction to successfully obtain key influencing factors, and use these factors to guide the production process, which improves production efficiency and reduces production commissioning costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179706A_ABST
    Figure CN120179706A_ABST
Patent Text Reader

Abstract

The invention discloses an extraction and mining application method for production data. The method comprises the following steps: S1, collecting the production data; s2, segmenting the production class data according to the data generation time, and integrating the production class data according to a specific index name to obtain segmented and integrated data; s3, performing data feature dimensionality reduction and clustering on the segmented and integrated data to obtain a plurality of clusters; extracting dimension-reduced data samples from the plurality of clustering clusters to obtain final dimension-reduced data samples extracted from all the clustering clusters; and S4, adopting a random forest algorithm to carry out importance ranking on the finally extracted dimension-reduced data samples of all the clusters, identifying key influence factors, and then adopting the key influence factors to guide a production process. According to the extraction and mining application method of the production data, the production data is deeply mined, the key influence factors are obtained to guide the production process, the production efficiency is improved, and the production debugging cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data mining in manufacturing, and particularly relates to a method for extracting, mining and applying production data. Background Art

[0002] Electronic manufacturing data has the characteristics of large capacity, multiple types, high dimension, high qualification rate and rapid update, and its main law is not obvious itself.

[0003] For data analysis and mining work, it is necessary to mine and analyze the information contained in a large amount of incomplete, noisy and random data, so as to extract key influencing factors (i.e., potential main laws) and useful information, and take corresponding measures to create value. Therefore, it is necessary to develop a method to achieve the purpose of extracting key influencing factors and useful information. Summary of the Invention

[0004] The purpose of the present invention is to propose a method for extracting, mining and applying production data in view of the above technical problems.

[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0006] The present invention proposes a method for extracting, mining and applying production data, including the following steps:

[0007] Step S1: Collect production data;

[0008] Step S2: Cut the production data according to the data generation time, and then integrate it according to a specific index name to obtain the cut and integrated data.

[0009] Step S3: After dimensionality reduction and clustering of the cut and integrated data, obtain multiple clustering clusters; then extract the dimensionality-reduced data samples from the multiple clustering clusters respectively to obtain the dimensionality-reduced data samples finally extracted by all the clustering clusters;

[0010] Step S4: Use the random forest algorithm to rank the importance of the dimensionality-reduced data samples finally extracted by all the clustering clusters, identify the key influencing factors, and then use the key influencing factors to guide the production process.

[0011] In the method for extracting, mining and applying production data of the present invention, the index name is the equipment name or the equipment number.

[0012] In the method for extracting, mining and applying production data of the present invention, the production data includes equipment output data, equipment status data and equipment alarm data.

[0013] In the method for extracting, mining and applying production data of the present invention, step S3 includes:

[0014] Step S3.1: Perform data feature dimensionality reduction on the segmented and integrated data according to the PCA method to obtain the dimensionality-reduced data. Among them, the dimensionality-reduced data contains n dimensionality-reduced data samples.

[0015] Step S3.2: Process the dimensionality-reduced data using the K-means clustering algorithm to obtain K clustering clusters and find the clustering centers of each clustering cluster.

[0016] Step S3.3: Statistically analyze the sample size of each clustering cluster.

[0017] The order of the clustering clusters is represented by i, and the sample size of the i-th clustering cluster is denoted as Ni. For the i-th clustering cluster, extract Ni / 1000 dimensionality-reduced data samples and calculate the variance of the extracted Ni / 1000 dimensionality-reduced data samples.

[0018] Step S3.5: For the i-th clustering cluster, traverse all extraction combinations of the Ni / 1000 dimensionality-reduced data samples and calculate the variance of the extracted Ni / 1000 dimensionality-reduced data samples respectively.

[0019] Among them, the variance calculated using the Ni / 1000 dimensionality-reduced data samples extracted in the j-th extraction is denoted as S i,j ;

[0020] Step S3.6: Calculate the moving adjacent difference δ i,j+1 = s i,j+1 - s i,j , thereby obtaining the array Δ i = [δ i,2 , δ i,3 ,..., δ i,1000 ;

[0021] Step S3.7: Obtain the order corresponding to the element with the maximum value in the array Δ i , denoted as p. Then, the number of dimensionality-reduced data samples ni = (p - 1) × Ni / 1000 finally extracted from the i-th clustering cluster and the corresponding sample data can be obtained.

[0022] Step S3.8: Traverse all the clustering clusters to obtain all the dimensionality-reduced data samples finally extracted from the clustering clusters and their total quantity.

[0023] In the above extraction and mining application method for production data of the present invention, the segmented and integrated data contains n segmented and integrated data samples; each segmented and integrated data sample contains m features.

[0024] Step S3.1 includes:

[0025] Form the segmented and integrated data into a data sample matrix A n×m ; Among them,

[0026]

[0027] x 11 …x 1m , respectively representing the 1st to the mth eigenvalue of the 1st segmented and integrated data sample x1;

[0028] …

[0029] x n1 …x nm , respectively representing the 1st to the mth eigenvalue of the nth segmented and integrated data sample x n ;

[0030] Center the data sample matrix A n×m to obtain the centered data sample matrix A' n×m ;

[0031] Calculate the covariance matrix B n×m of the centered data sample matrix A' m×m , and perform eigen-decomposition on the covariance matrix B m×m to obtain multiple eigenvalues and their corresponding multiple eigenvectors;

[0032] Then, sort the multiple eigenvalues from largest to smallest, select the first j eigenvalues in sequence from the beginning, and select the corresponding eigenvectors, and output the projection matrix C composed of the selected eigenvectors m×j ;

[0033] By projecting the centered data sample matrix A' n×m onto C m×j to obtain the dimensionality-reduced data sample matrix D n×j , that is, obtain the dimensionality-reduced data; where D n×j = A' n×m * C m×j .

[0034] The extraction and mining application method of production data of the present invention deeply mines the production data, adopts the methods of segmentation, dimensionality reduction, and extraction to obtain the key influencing factors, and uses the key influencing factors to guide the production process, improving the production efficiency and reducing the production debugging cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0036] Figure 1 Shows the flow schematic diagram of the extraction and mining application method of production data of the present invention.

[0037] Figure 2The flowchart of step S2 of the extraction, mining and application method for production - related data of the present invention is shown. Detailed implementation manners

[0038] As Figure 1 shown, the present invention provides an extraction, mining and application method for production - related data, including the following steps:

[0039] Step S1: Collect production - related data;

[0040] In this step, the production - related data includes MES system data (i.e., manufacturing enterprise production process execution system data), manual record data, sensor data, and equipment management system data.

[0041] Step S2: Split the production - related data according to the data generation time, and then integrate it according to a specific index name to obtain the split - integrated data, as Figure 2 shown.

[0042] In this step, the index name can be the equipment name or the equipment number.

[0043] Here, by splitting the production - related data according to the data generation time, production - related data at different times can be obtained, including equipment output data, equipment status data, and equipment alarm data.

[0044] Through the operation of splitting according to the data generation time, it is possible to automatically identify the production - related data in a certain time period, so as to batch - process multi - dimensional data. For example, obtain the status data of multiple devices at one time, and then batch - summarize according to the data generation time to achieve statistics.

[0045] Step S3: After performing data feature dimensionality reduction and clustering on the split - integrated data, obtain multiple clustering clusters; then extract the dimensionality - reduced data samples from each of the multiple clustering clusters to obtain all the dimensionality - reduced data samples finally extracted from the clustering clusters;

[0046] Step S3 includes:

[0047] Step S3.1: Perform data feature dimensionality reduction on the split - integrated data according to the PCA method (Principal Component Analysis), to obtain the dimensionality - reduced data;

[0048] Here, the PCA method transforms the data into a new coordinate system through a linear transformation, that is, converts multiple indicators into a few comprehensive indicators and retains the features that contribute the most to the variance of the data set.

[0049] Specifically, the split - integrated data contains n split - integrated data samples; each split - integrated data sample contains m features; step S3.1 includes:

[0050] Form the data sample matrix A by integrating the segmented data n×m ; where

[0051]

[0052] x 11 …x 1m , representing the 1st - mth eigenvalues of the 1st segmented and integrated data sample x1 respectively;

[0053] …

[0054] x n1 …x nm , representing the 1st - mth eigenvalues of the nth segmented and integrated data sample x n respectively;

[0055] Center the data sample matrix A n×m to obtain the centered data sample matrix A' n×m ;

[0056] Calculate the covariance matrix B of the centered data sample matrix A' n×m , and perform eigen - decomposition on the covariance matrix B m×m to obtain multiple eigenvalues and their corresponding multiple eigen - vectors; m×m Then, sort the multiple eigenvalues from large to small, select j eigenvalues in sequence starting from the beginning, and select the corresponding eigen - vectors, and output the projection matrix C composed of the selected eigen - vectors

[0057] ; m×j ;

[0058] By projecting the centered data sample matrix A' n×m onto C m×j obtain the dimensionality - reduced data sample matrix D n×j , that is, obtain the dimensionality - reduced data; where D n×j = A' n×m *C m×j ; where the dimensionality - reduced data contains n dimensionality - reduced data samples.

[0059] In step S3.1, use the formula B * ζ = λ * ζ to obtain the eigenvalue λ of the covariance matrix B and its corresponding eigen - vector ζ.

[0060] Step S3.2: Use the K - means clustering algorithm to process the dimensionality - reduced data, obtain K clustering clusters, and find the cluster center of each cluster;

[0061] In step S3.2, the number of clusters K is selected according to the sum of squared errors of sample clustering Proceed. The larger the K, the smaller the SSE, indicating a higher degree of aggregation among samples. When K is less than the true number of clusters, as K increases, the degree of aggregation of each cluster will increase, and the decline rate of SSE will be large. When K reaches the true number of clusters, if K continues to increase, the degree of aggregation obtained will rapidly decrease, the decline rate of SSE will suddenly decrease and tend to level off as K continues to increase. The point that first levels off is the appropriate number of clusters K. Then, calculate the Euclidean distance between data. For example, the formula example in a two-dimensional space is Finally, find the cluster centers; finally, count the labels of each cluster and the corresponding number of samples. Starting from the cluster center of each cluster, sort the corresponding labeled data in ascending order of distance and number the data with natural numbers such as 1, 2, 3…n.

[0062] Step S3.3: Count the sample size of each cluster;

[0063] Step S3.4: Use i to represent the order of the cluster. Denote the sample size of the i-th cluster as Ni; for the i-th cluster, extract Ni / 1000 dimensionality-reduced data samples and calculate the variance of the Ni / 1000 dimensionality-reduced data samples extracted;

[0064] Specifically, in Step S3.4, each time the number of extracted samples is Ni / 1000, and calculate the variance s of the extracted sample data i,j Perform a loop. After each loop ends, update the starting position number to bi = bi + Ni / 1000. The total number of loops is 1000, and obtain the variance array S of all iterations of the corresponding cluster i,j =[s i,1 ,s i,2 ,...,s i,1000 ;

[0065] Step S3.5: For the i-th cluster, traverse all extraction combinations of Ni / 1000 dimensionality-reduced data samples, and calculate the variance of the Ni / 1000 dimensionality-reduced data samples extracted respectively;

[0066] Among them, use the Ni / 1000 dimensionality-reduced data samples extracted in the j-th extraction, and denote the calculated variance as S i,j ;

[0067] Step S3.6: Calculate the difference δ between adjacent movements i,j+1 =s i,j+1 -s i,j , so as to obtain the array Δ i =[δ i,2 ,δ i,3 ,...,δ i,1000 ;

[0068] Step S3.7: Obtain the array Δi The order corresponding to the element with the maximum value among them is denoted as p; thus, the number of finally extracted and dimension-reduced data samples ni for the i-th clustering cluster can be obtained as ni = (p - 1) × Ni / 1000, along with the corresponding sample data.

[0069] Step S3.8: Traverse all the clustering clusters to obtain all the finally extracted and dimension-reduced data samples of all the clustering clusters and their total quantity.

[0070] In this step S3.8, the total quantity of all the finally extracted and dimension-reduced data samples of all the clustering clusters can be denoted as:

[0071]

[0072] Step S4: Use the random forest algorithm to rank the importance of all the finally extracted and dimension-reduced data samples of all the clustering clusters, identify the key influencing factors, and then use the key influencing factors to guide the production process.

[0073] In this step, the key parameters used in the random forest algorithm include: the number of decision trees is 100, the maximum depth is 5, and the minimum number of samples in a leaf node is 5.

[0074] In order to make the technical objectives, technical solutions, and technical effects of the present invention clearer, so as to facilitate those skilled in the art to understand and implement the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0075] The present invention proposes a method for extracting, mining and applying production - type data, including the following steps: Step S1: Data collection; Step S2: Data processing and integration; Step S3: Construction of an adaptive data extraction method; Step S4: Data mining analysis and application. When implementing this embodiment, in the first step, production - type data is collected first, including through - station production data, equipment status data, and equipment alarm data. In the second step, the collected production - type data is processed and integrated. Step S21: Taking the equipment as a loop, each equipment has state types such as running, waiting for materials, and downtime, as well as the alarm record status data of each equipment. A module for segmenting and reorganizing data by time period is established. The function of this module can segment the equipment production data, equipment status data, and alarm record data according to the required time - period granularity and then reorganize them with the equipment as the index, obtaining data such as the duration, production volume, and number of alarms of each state type of each equipment under each time granularity. Step S22: Through the equipment loop, for each moment of each state type of all equipment, the durations are added correspondingly with the equipment and shift (or moment) as the abscissa and the equipment status as the ordinate. In this way, data such as production volume data, corresponding durations of operating states (such as running, waiting for materials, stopping, etc.), number of alarms, and alarm durations under the shift (or moment) of the equipment are obtained. Step S23: Calculate the equipment efficiency achievement rate. The efficiency achievement rate = actual output / theoretical output. The actual output is the sum of all hourly production volumes of each shift, and the theoretical output is the theoretical hourly production capacity of the equipment for producing the current product model multiplied by the actual equipment operation duration. Step S24: Integrate with the equipment column as the index, obtaining data with the equipment efficiency achievement rate as the target attribute and the production volume data, operating state duration data, and alarm state data of each equipment shift as the feature attributes. In the third step, an adaptive data extraction method is constructed. The PCA (Principal Component Analysis) method is used to reduce the dimension of the data integrated in the second step to three - dimensional, and then clustering analysis is performed. The K - Means clustering analysis method is adopted, and the number of clusters K value is selected as 2, obtaining two clusters of data. According to the adaptive data extraction method, a data sample of 1120 rows is finally extracted. In the fourth step, data mining analysis and application are carried out. The sample data extracted in the third step is deeply mined and analyzed, and the importance of feature attribute factors is ranked using the random forest algorithm. Among them, the key parameters used in the random forest algorithm include: the number of decision trees is 100, the maximum depth is 5, and the minimum number of samples in the leaf node is 5. The key influence feature attribute factors on the target attribute are identified, which are its potential main rules. Finally, this discovery point is used to guide the improvement of the process flow. After the improvement action, the equipment efficiency achievement rate has increased by about 16%.

[0076] It should be understood that for those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for extracting and mining application of production - type data, characterized in that, It includes the following steps: Step S1, collect production data; Step S2, split the production data according to the data generation time, and then integrate it according to a specific index name to obtain the split and integrated data; Step S3, after performing data feature dimensionality reduction and clustering on the split and integrated data, obtain multiple clustering clusters; then extract the dimensionality-reduced data samples from each of the multiple clustering clusters respectively to obtain the dimensionality-reduced data samples finally extracted by all the clustering clusters; Step S4, use the random forest algorithm to rank the importance of the dimensionality-reduced data samples finally extracted by all the clustering clusters, identify the key influencing factors, and then use the key influencing factors to guide the production process.

2. The method for extracting and mining application of production - type data according to claim 1, characterized in that, The index name is the equipment name or the equipment number.

3. The method for extracting and mining application of production - type data according to claim 2, characterized in that, The production data includes equipment output data, equipment status data, and equipment alarm data.

4. The method for extracting and mining application of production - type data according to claim 1, characterized in that, Step S3 includes: Step S3.1, perform data feature dimensionality reduction on the split and integrated data according to the PCA method to obtain the dimensionality-reduced data; among them, the dimensionality-reduced data contains n dimensionality-reduced data samples; Step S3.2, use the K-means clustering algorithm to process the dimensionality-reduced data to obtain K clustering clusters, and find the clustering center of each clustering cluster; Step S3.3, count the sample size of each clustering cluster; The order of the clustering clusters is represented by i, and the sample size of the i-th clustering cluster is denoted as Ni; for the i-th clustering cluster, extract Ni / 1000 dimensionality-reduced data samples, and calculate the variance of the Ni / 1000 dimensionality-reduced data samples extracted; Step S3.5, for the i-th clustering cluster, traverse all extraction combinations of the Ni / 1000 dimensionality-reduced data samples, and calculate the variance of the Ni / 1000 dimensionality-reduced data samples extracted respectively; Among them, the variance calculated using Ni / 1000 dimensionality-reduced data samples extracted in the j-th extraction is denoted as S i,j ; Step S3.6: Calculate the difference δ between adjacent movements i,j+1 = s i,j+1 - s i,j to obtain the array Δ i = [δ i,2 , δ i,3 ,..., δ i,1000 ; Step S3.7: Obtain the array Δ i The order corresponding to the element with the maximum value in i is denoted as p. Then, the number of the finally extracted dimensionality-reduced data samples of the i-th clustering cluster, ni = (p - 1) × Ni / 1000, and the corresponding sample data can be obtained. Step S3.8, traverse all the clustering clusters to obtain the dimensionality-reduced data samples finally extracted by all the clustering clusters and their total quantity.

5. The method for extracting and mining application of production - type data according to claim 4, characterized in that, The split and integrated data contains n split and integrated data samples; each split and integrated data sample contains m features; Step S3.1 includes: Form the segmented and integrated data into a data sample matrix A n×m ; where x 11 …x 1m , respectively representing the first to the m-th eigenvalue of the first segmented and integrated data sample x1; … x n1 …x nm represent the first to the m-th eigenvalue of the n-th segmented and integrated data sample x n respectively; Center the data sample matrix A n×m to obtain the centered data sample matrix A' n×m ; Calculate the covariance matrix B of the centralized data sample matrix A' n×m and perform eigen-decomposition on the covariance matrix B m×m to obtain multiple eigenvalues and their corresponding multiple eigenvectors; m×m ​ Then, sort the multiple eigenvalues from largest to smallest, sequentially select j eigenvalues starting from the beginning, and select the corresponding eigenvectors, and output the projection matrix C composed of the selected eigenvectors m×j ; By projecting the centralized data sample matrix A' n×m onto C m×j the dimensionality-reduced data sample matrix D is obtained n×j , that is, the dimensionality-reduced data is obtained; where D n×j = A' n×m * C m×j .