High-Dimensional Feature Reconstruction and Fusion Method Based on SMT Quality Big Data

By extracting and fusing text data and structured data from SMT production lines, a SAE feature reconstruction model with different activation functions is constructed, the features are iteratively trained and weighted fusion are generated, and high-dimensional fusion features are generated, which solves the problems of low data utilization and poor generalization capabilities in the existing technology, and achieves higher prediction accuracy and data utilization.

CN116628623BActive Publication Date: 2025-07-01XIDIAN UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310590118.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-07-01
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

The existing SMT production line intelligent decision-making algorithm mainly relies on structured data, fails to make full use of text data, and lacks feature dimensions, resulting in low model generalization capabilities and prediction accuracy.

Method used

A high-dimensional feature reconstruction and fusion method based on SMT quality big data is proposed. By acquiring and preprocessing text data and structured data, extracting and merging features, building a stacked autoencoder (SAE) feature reconstruction model with different activation functions, iteratively training and weighting fusion of features, and generating high-dimensional fusion features.

Benefits of technology

The data utilization rate is improved, the generalization ability and prediction accuracy of SMT quality prediction models are improved, and the problems of insufficient feature dimensions and poor generalization ability of the model are solved in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116628623B_ABST
    Figure CN116628623B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-dimensional feature reconstruction and fusion method based on SMT quality big data, which mainly solves the problems of low data utilization, too few feature dimensions and low accuracy of prediction models in the prior art. Its implementation scheme is: preprocessing the SMT production line text data set and structured data set; constructing and training a text data feature extraction model to obtain text data extraction features; constructing a structured data feature extraction model to obtain structured data extraction features; merging and deduplicating features extracted from text data and structured data; using a stacked autoencoder and a method based on the combination of mean square error and mean absolute percentage error to reconstruct and fuse the extracted features. The present invention improves the utilization rate of SMT enterprise data, realizes the fusion of text data and structured data, increases the data dimension to more than 50 dimensions, improves the accuracy of the model, and can be used for multimodal data processing of SMT production line quality big data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of physical technologies, and further relates to a high-dimensional feature reconstruction and fusion method, which can be used for multi-modal data processing of quality big data in the SMT production line. Technical Background

[0002] Electronic manufacturing enterprises have accumulated a large amount of text data and structured data during production, but they all exist in isolation. At present, most of the intelligent decision-making algorithms in the SMT production line only utilize structured data and have not considered integrating text data into the algorithms. The potential value therein is mostly ignored, and the dimension of the current features is mostly below 20 dimensions, without considering the important role of high-dimensional non-linear features in algorithms. Fully exploring the value of text data and realizing the reconstruction and fusion of text data and structured data are of great significance for improving the SMT production line process and product quality. The use of entity recognition, entity extraction technology, and graph convolutional neural network technology can achieve feature extraction of text data, and the use of data mining technology can achieve feature extraction of structured data. The stacked autoencoder realizes the feature fusion of text data and structured data.

[0003] Jiangsu Dake Digital Technology Co., Ltd. disclosed a "data processing method and platform applicable to system security operation and maintenance" in its patent document with the application number 202310045129.3. The implementation steps are as follows: First step, data processing is carried out using the principle of texture primitive histogram; Second step, feature extraction of image data is carried out based on a convolutional neural network; Third step, feature extraction of text data is carried out based on a recurrent neural network; Fourth step, feature reconstruction is carried out on at least one feature related to production equipment failures to obtain a reconstructed feature set. Since this method only utilizes the text data generated from images during the production process and does not consider reconstructing and fusing structured data and text data, the model is not fully trained, resulting in weak generalization ability and low data utilization rate.

[0004] Chengdu Anze Technology Co., Ltd. discloses "A Method, System, Terminal and Medium for Identifying Radio Frequency Hopping Signals" in its patent document with the application number 202310000571.4. The implementation steps are as follows: First, obtain the radio frequency hopping signal of the target source, and preprocess the radio frequency hopping signal to obtain time-domain information; Second, perform short-time Fourier transform on the time-domain signal, and dynamically adjust the window width of the sliding window according to the waveform amplitude distribution characteristics and waveform time distribution characteristics of the time-domain signal to obtain the frequency-domain signal; Third, use the time-frequency distribution method to process the frequency-domain signal and the time-domain signal to obtain the time-frequency feature map; Fourth, extract the time-frequency features of a single type and the correlation features between the time-frequency features in the time-frequency feature map, and reconstruct the recognition features based on the time-frequency features and the correlation features; Fifth, input the recognition features into the pre-constructed neural network recognition model for training and recognition to obtain the recognition result of the radio frequency hopping signal. Since the reconstructed features are only below 20 dimensions, this method fails to consider the non-linear relationship between high-dimensional features and the knowledge in the text data, and fails to fuse structured data and text data, resulting in a low prediction accuracy of the model. Summary of the Invention

[0005] The purpose of the present invention is to propose a high-dimensional feature reconstruction and fusion method based on SMT quality big data in view of the above-mentioned deficiencies of the prior art, so as to improve the utilization rate of SMT quality big data and enhance the generalization ability and prediction accuracy of the SMT quality prediction model.

[0006] The technical solutions for realizing the purpose of the present invention include the following steps:

[0007] (1) Obtain the text data and structured data in the SMT production line quality big data, and preprocess the text data and structured data respectively to obtain the preprocessed SMT production line text data set and SMT production line structured data set;

[0008] (2) Extract the features of the preprocessed SMT production line pre-text data set and SMT production line structured data set respectively to obtain the text feature set and data feature set;

[0009] (3) Merge and deduplicate the text feature set and the data feature set, and screen the original structured data set according to the merged and deduplicated feature set, and divide the screened data set into four categories: process, quality, production, equipment, and others except these four categories;

[0010] (4) Construct 5 integrated stack autoencoder SAE feature reconstruction models with different activation functions, including encoders and decoders:

[0011] Establish an SAE feature reconstruction model with Tanh as the activation function;

[0012] Build an SAE feature reconstruction model with the Sigmoid activation function;

[0013] Build an SAE feature reconstruction model with the Relu activation function;

[0014] Build an SAE feature reconstruction model with the Softmax activation function;

[0015] Build an SAE feature reconstruction model with the ReLU6 activation function;

[0016] (5) Iteratively train the SAE feature reconstruction model constructed in (4) to obtain the feature reconstruction result:

[0017] (5a) Input the process, quality, production, equipment, and other datasets except these four categories into 5 different SAE feature reconstruction models, and output the predicted values of the quality indicators obtained for each category of data in each SAE feature reconstruction model;

[0018] (5b) Set the loss function MSE for each SAE feature reconstruction model, and calculate the loss value J of the SAE feature reconstruction model through the predicted values and actual values of the quality indicators k ;

[0019] (5c) Calculate the loss gradient of the loss value J k by the backpropagation method Then adopt the stochastic gradient descent method, and through the loss gradient update the weights w k of the encoder and decoder in the SAE feature reconstruction model until J k < 0.1, then stop training, and take the output result of the last iteration as the feature reconstruction result of each category of data;

[0020] (6) Fuse the feature reconstruction results of each category of data:

[0021] (6a) Take the loss value J k obtained in step (5b) as the mean square error M of each SAE feature reconstruction model in each category of data, and based on the predicted values of the quality indicators corresponding to each category of data in step (5a), obtain the mean absolute percentage error MAPE of each SAE feature reconstruction model in each category;

[0022] (6b) Based on the MSE and MAPE of the SAE feature reconstruction model, perform weighted fusion on the dataset reconstructed each time to obtain the fusion result:

[0023] (6b1) Calculate the collaborative error of each category of SAE feature reconstruction model:

[0024]

[0025] Among them, E j represents the collaborative error of the j-th SAE feature reconstruction model in each type of dataset based on the j-th mean square error M and the j-th mean absolute percentage error MAPE, where j is an integer in [1, 5];

[0026] (6b2) Calculate the weight of each type of SAE feature reconstruction model in each type of dataset according to the collaborative error:

[0027]

[0028] Among them, w j is the weight of each type of dataset in the j-th SAE feature reconstruction model, and n is the number of SAE feature reconstruction models corresponding to each type of dataset;

[0029] (6b3) According to the weights of each type of SAE feature reconstruction model in each type of dataset, fuse the feature reconstruction results of each type of data to obtain the new feature F of each type of dataset a :

[0030]

[0031] Among them, x is the feature reconstruction result of the data of the current category, and a is an integer in [1, 5];

[0032] (6c) Merge the new features F a of each type of dataset, that is, splice the corresponding columns of the features to obtain the fusion result F.

[0033] The present invention has the following advantages compared with the prior art:

[0034] First, since the present invention extracts the text data and structured data features, merges and de-duplicates the two types of features, and combines and uses the text data and structured data, it solves the deficiency of only using single-type data in the prior art and improves the utilization rate of data.

[0035] Second, since the present invention establishes an SAE feature reconstruction model based on 5 different activation functions to reconstruct data in the SMT data reconstruction, it solves the problem of only using a single activation function to construct a reconstruction model in the prior art and further improves the utilization rate of data.

[0036] Thirdly, in the SMT data processing of the present invention, first, the collaborative error of each type of SAE feature reconstruction model is calculated, then the weight of each type of SAE feature reconstruction model in each type of dataset is calculated using the collaborative error, and finally, the data of each category is weighted and summed to form a new fused feature set, which solves the problem that the existing technology fusion method only performs merging, resulting in too low feature dimensions and low accuracy of the prediction model, and improves the generalization ability and prediction accuracy of the SMT quality prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is the implementation flowchart of the present invention;

[0038] Figure 2 is the annotation schematic diagram of entities and entity relationships in the present invention;

[0039] Figure 3 is the sub-flowchart of extracting SMT production line text data features in the present invention;

[0040] Figure 4 is the sub-flowchart of extracting SMT production line structured data features in the present invention;

[0041] Figure 5 is the schematic diagram of reconstructing and fusing SMT production line features in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0043] This example is based on the quality big data in the SMT production line for reconstruction and fusion. The SMT production line is a typical production line in the electronic manufacturing industry. The production and manufacturing of electronic products on the market are inseparable from the SMT production line. Nowadays, the electronic manufacturing industry is not only an important part of the national economy, but also an important indicator that enables a country to stand firm on the world stage and cannot be ignored. The purpose of the present invention is to improve the data utilization rate of the SMT production line, and enhance the generalization ability and prediction accuracy of the SMT quality prediction model.

[0044] Refer to Figure 1 , the implementation steps of this embodiment are as follows:

[0045] Step 1: Obtain the text data and structured data in the SMT production line quality big data, and preprocess the text data and structured data respectively to obtain the preprocessed SMT production line text dataset and SMT production line structured dataset.

[0046] 1.1) Collect the SMT production line text dataset, which includes a developer's manual, 100 process documents, 200 quality feedback forms, and 10 design documents as the knowledge sources of the text data;

[0047] 1.2) Preprocess the collected SMT production line text dataset:

[0048] 1.2.1) Delete the irrelevant data in each text dataset, remove the serial numbers and spaces that are irrelevant to the key information, fill in the missing values in the text data in combination with business knowledge, and cut the text according to the nearest full stop or exclamation mark to obtain the preliminarily processed text dataset;

[0049] 1.2.2) Based on data mining technology, take the product defect types, defect causes, solutions, defect phenomena, influencing factors, and defect consequences of the SMT production line as the knowledge ontology;

[0050] 1.2.3) According to the knowledge ontology, use YEDDA software to perform entity annotation on the preliminarily processed text dataset. The annotation instructions are shown in Table 1:

[0051] Table 1 Annotation instructions for the knowledge ontology

[0052]

[0053] 1.2.4) According to the BIO sequence annotation method, perform BIO sequence annotation on the preliminarily processed text dataset, that is, mark the first character of each entity sequence as "B - entity name"; mark the middle characters as "I - entity name"; mark the irrelevant characters as "O" to obtain the text dataset after entity annotation;

[0054] 1.2.5) Annotate the relationships between all pairs of entities in the text dataset after entity annotation to obtain the preprocessed text dataset.

[0055] In the embodiments of the present invention, the relationship names between pairs of entities and the corresponding annotation instructions are shown in Table 2:

[0056] Table 2 Entity relationship annotation instructions

[0057]

[0058] In the embodiments of the present invention, the relationship names between pairs of entities and the corresponding annotation schematic diagrams are as Figure 2 shown.

[0059] Figure 2In it, "too low viscosity" is the cause of the defect, so this entity is marked with R. Among them, "viscosity" is the first character of the entity sequence, marked as B-R; "degree", "excessive", and "low" are the middle characters, marked as I-R; "collapse" is the defect type, so this entity is marked with D. Among them, "collapse" is the first character, marked as B-D; "subsidence" is the middle character, marked as I-D; the remaining words are irrelevant character marks, marked as O; the entity relationship between "too low viscosity" and "collapse" is that "collapse" is caused by "too low viscosity", marked as RCD.

[0060] 1.3) Preprocess the collected structured dataset of the SMT production line:

[0061] In this example, the structured dataset of the SMT production line of a certain research institute of CETC is used as the data source. This SMT structured dataset contains nearly ten million production data of the company in the past year. The data are all csv structured data, and the content is shown in Table 3. Among them, the features of the dataset are squeegee pressure, squeegee speed, printing height compensation, table separation speed, automatic cleaning count, cleaning speed, table separation distance, cleaning supply time, and squeegee separation distance; the quality indicators of the dataset are volume, area, height, X offset, and Y offset.

[0062] Table 3 List of parameters of the structured dataset of the SMT production line

[0063]

[0064] The preprocessing steps for this dataset are as follows:

[0065] 1.3.1) Through statistical analysis, the mode of the squeegee pressure field is 12, so use 12 to fill the missing value NaN of the squeegee pressure in the 3rd data;

[0066] 1.3.2) Through statistical analysis, the mode of the squeegee speed field is 20, so use 20 to fill the missing value NaN of the squeegee speed in the 4th data;

[0067] 1.3.3) Through statistical analysis, the mode of the table separation speed is 0.333, so use 0.333 to fill the missing values NaN of the table separation speed in the 5th and 6th data;

[0068] 1.3.4) Use the normal distribution and box plot to detect outliers in all data, and detect that the volume of the 6th data is an outlier, so delete the 6th data;

[0069] 1.3.5) Use the z-score method to standardize the original data to make it present a normal distribution. The formula is:

[0070]

[0071] Among them, x * is the standardized value, x is the value of the corresponding column of the dataset feature, u is the mean of the column to which x belongs, and σ is the standard deviation of the column to which x belongs.

[0072] After all the above processing, the preprocessed structured dataset of the SMT production line is obtained, as shown in Table 4:

[0073] Table 4 Results of the preprocessing of the structured data of the SMT production line

[0074]

[0075] Step 2: Extract the features of the preprocessed text dataset of the SMT production line to obtain a text feature set.

[0076] Existing methods for extracting text dataset features include the Boolean model, the vector space model, the graph space model, the BERT-Bi-LSTM-CRF named entity recognition model, the BERT entity relation extraction model, or a combination of several models. In this example, but not limited to, a combination of the BERT-Bi-LSTM-CRF named entity recognition model, the BERT entity relation extraction model, and the two-layer graph convolutional neural network GCN model is used.

[0077] Refer to Figure 3 , and the specific implementation of this step is as follows:

[0078] 2.1) Construct a BERT-Bi-LSTM-CRF named entity recognition model composed of a series connection of a BERT embedding layer, a Bi-LSTM layer, and a CRF layer, where:

[0079] For the BERT embedding layer, the number of network layers is set to 10, the number of hidden units is set to 384, and the number of attention heads is set to 10;

[0080] For the Bi-LSTM layer, the parameters of each neuron are initialized using the Xavier method;

[0081] For the CRF layer, the output result is initialized using the randn function.

[0082] 2.2) Train the BERT-Bi-LSTM-CRF named entity recognition model:

[0083] Input the preprocessed pre-text dataset of the SMT production line into the BERT-Bi-LSTM-CRF named entity recognition model to obtain the annotation sequence of all named entities;

[0084] Using the mean square error formula MSE, calculate the loss value between the predicted annotation sequence and the actual annotation sequence of each feature vector sequence. For each feature vector sequence's loss value, according to the stochastic gradient descent method, based on the loss values of all feature vector sequences, backpropagate to adjust the number of neurons in the Bi-LSTM layer until the loss value is less than or equal to 0.1. Take the output result of the last iteration as the named entity recognition result.

[0085] The mean square error formula is expressed as follows:

[0086]

[0087] where n represents the number of pre-text datasets of the SMT production line, y i represents the actual annotation sequence of each feature vector, represents the predicted annotation sequence of each feature vector.

[0088] 2.3) Set the initial parameters of the existing BERT entity relation extraction model and train it:

[0089] 2.3.1) Set the maximum number of words to 64, the batch data size to 64, the learning rate to 1×10 -5 , the dropout rate to 0.3, and the number of iterations to 10 times;

[0090] 2.3.2) Input each feature vector sequence in the named entity recognition result into the BERT entity relation extraction model to obtain the relation vector between pairwise named entities;

[0091] 2.3.2) Using the same error formula as in step 2.2), calculate the loss value between the predicted relation annotation sequence and the actual relation annotation sequence of each feature vector sequence. For each feature vector sequence's loss value, according to the stochastic gradient descent method, based on the loss values of all feature vector relation sequences, adjust the learning rate and the dropout rate until the loss value is less than or equal to 0.1. Take the output result of the last iteration as the knowledge extracted from the text data.

[0092] 2.4) Represent the knowledge extracted from the text data in the form of triples, that is, store each entity in the form of triples <entity, attribute name, attribute value>, establish the relation connection between entities, and ensure the consistency of data description between the two connected entities, ensuring that the triples satisfy the form <entity 1, relation, entity 2>.

[0093] In this embodiment, the triples are represented as:

[0094] <Adjust the stencil opening, avoid, bridging>

[0095] <The stencil opening is too large, resulting in bridging>

[0096] Among them, the meaning represented by the first triple is that adjusting the stencil opening can avoid bridging defects, and the meaning represented by the second triple is that too large stencil opening will cause bridging defects.

[0097] 2.5) Store knowledge based on the Neo4j graph relational database, import the knowledge in the form of triples into the Neo4j graph relational database, and form the SMT production line quality knowledge graph.

[0098] 2.6) Use the existing Glove-word-vector word vector representation method to represent the knowledge graph in the form of word vectors.

[0099] 2.7) Select the existing two-layer graph convolutional neural network GCN model and train it:

[0100] 2.7.1) Set the initialization parameters of the two-layer graph convolutional neural network GCN model: set the number of neurons in the first layer of GCN to 16, the number of neurons in the second layer of GCN to 7, the loss function to MSE, the learning rate to 0.1, and the number of iterations to 200;

[0101] 2.7.2) Input the word vectors into the two-layer graph convolutional neural network GCN model to obtain the word vector prediction sequence;

[0102] 2.7.3) Use the same error formula as in step 2.2) to calculate the loss value between each predicted word vector sequence and the actual word vector sequence, adjust the learning rate according to the loss value of each word vector sequence by the stochastic gradient descent method until the loss value is less than or equal to 0.1, and obtain the weights of the model in the last iteration.

[0103] 2.8) Sort the obtained weights, and take the features with weights greater than or equal to 0.5 as the features extracted from the text data set to obtain the text feature set.

[0104] Step 3, use the existing XGBoost feature extraction model to extract the features of the preprocessed SMT production line structured data set to obtain the data feature set.

[0105] Refer to Figure 4 , The specific implementation of this step is as follows:

[0106] 3.1) Set the initialization parameters of the XGBoost feature extraction model as shown in Table 5:

[0107] Table 5 Initialization information table of key parameters of the XGBoost model

[0108]

[0109] 3.2) Input the features in the preprocessed structured data set of the SMT production line into the XGBoost feature extraction model, and output the predicted values of the data set quality indicators respectively;

[0110] 3.3) Use the importance formula in the XGBoost model to calculate the importance of the influencing factors of each feature in the data set:

[0111]

[0112] where score i represents the importance of the influencing factors of the i-th feature in the data set, G L represents the sum of the first-order derivatives of all left leaf nodes in the XGBoost model, G R represents the sum of the first-order derivatives of all right leaf nodes in the XGBoost model, H L represents the sum of the second-order derivatives of all left leaf nodes in the XGBoost model, H R represents the sum of the second-order derivatives of all right leaf nodes in the XGBoost model, ρ and γ represent the regularization parameters when the loss function of the XGBoost model reaches the minimum;

[0113] 3.4) Statistically analyze the importance of the influencing factors of each feature, as shown in Table 6:

[0114] Table 6 Importance of the influencing factors of each feature

[0115]

[0116] 3.5) Merge the features with the maximum influencing factor importance greater than 100 into the data feature set.

[0117] Step 4, divide the data set according to the text feature set and the data feature set.

[0118] 4.1) Merge and deduplicate the text feature set and the data feature set to obtain a new feature set. Some features of the new feature set in this embodiment are shown in Table 7:

[0119] Table 7 New feature set after merging and deduplication

[0120]

[0121]

[0122] 4.2) Screen the original structured data set according to the new feature set, that is, extract and merge the columns with the same features as the original structured data set in the new feature set to form the screened structured data set, as shown in Table 8:

[0123] Table 8 Screened structured data set

[0124]

[0125] 4.3) Divide the filtered dataset into process, quality, production, equipment, and other datasets other than these four categories according to the data mechanism knowledge of the SMT production line and the logical relationship of data fields, that is, a total of 5 categories of datasets.

[0126] Step 5, construct 5 SAE feature reconstruction models with different activation functions.

[0127] Each model is composed of an encoder and a decoder connected in series, where:

[0128] The encoder has two layers. The number of hidden neurons in the first input layer is set to 20 according to the number of merged and processed features, and the number of hidden neurons in the second hidden layer is 10;

[0129] The decoder has two layers. The number of hidden neurons in the first hidden layer is 10, and the number of hidden neurons in the second output layer is 5;

[0130] The 5 activation functions are: hyperbolic tangent function Tanh, sigmoid function Sigmoid, rectified linear unit function Relu, normalized exponential function Softmax, ReLU6 function;

[0131] The 5 SAE feature reconstruction models constructed are as follows:

[0132] The first one: the SAE feature reconstruction model with the hyperbolic tangent function Tanh as the activation function ;

[0133] The second one: the SAE feature reconstruction model with the sigmoid function Sigmoid as the activation function ;

[0134] The third one: the SAE feature reconstruction model with the rectified linear unit function Relu as the activation function ;

[0135] The fourth one: the SAE feature reconstruction model with the normalized exponential function Softmax as the activation function ;

[0136] The fifth one: the SAE feature reconstruction model with the ReLU6 function as the activation function ;

[0137] Among them, x is the input neuron node value, and n is the number of input neurons.

[0138] Step 6, perform iterative training on the SAE feature reconstruction model constructed in step 4 to obtain the feature reconstruction result.

[0139] 6.1) Input process, quality, production, equipment, and other data sets other than these four categories into 5 different SAE feature reconstruction models, and output the predicted values of the quality indicators obtained for each category of data in each SAE feature reconstruction model. Since the five categories of data sets are input into 5 different SAE feature reconstruction models, a total of 25 predicted values of quality indicators are obtained;

[0140] 6.2) Set the loss function MSE of each SAE feature reconstruction model to be the same, which is expressed as follows:

[0141]

[0142] where N is the number of data sets in each category, represents the predicted value of the quality indicator of each category of data sets, and y i represents the actual value of the quality indicator of each category of data sets.

[0143] 6.3) According to the predicted values and actual values of the quality indicators, through the loss function set in step 6.2), calculate 25 loss values J k for the 5 SAE feature reconstruction models, where k ∈ [1, 25] and k is an integer;

[0144] 6.4) Calculate the loss gradient k of the loss value J

[0145]

[0146] where w k is the weight of the kth encoder and decoder;

[0147] 6.5) Adopt the stochastic gradient descent method to update the weights w of the encoder and decoder in the SAE feature reconstruction model through the loss gradient k , until J k < 0.1, then stop training, and take the output result of the last iteration as the feature reconstruction result of each category of data.

[0148] The update formula for the weight w k is as follows;

[0149]

[0150] where w k ' represents the updated result of w k , and α represents the learning rate, where α ∈ [0, 1].

[0151] Step 7, fuse the feature reconstruction results of each category of data.

[0152] Referring to Figure 5 , the specific implementation of this step is as follows:

[0153] 7.1) Take the loss value J k obtained in step 6.3) as the mean square error M of each SAE feature reconstruction model in each category of data;

[0154] 7.2) Calculate the mean absolute percentage error MAPE of each SAE feature reconstruction model in each category of data according to the quality index prediction values output by each SAE feature reconstruction model corresponding to each category of data in step 6.1):

[0155]

[0156] where N is the number of each category of data sets, represents the predicted value of the quality index of each category of data sets, y i represents the actual value of the quality index of each category of data sets;

[0157] 7.3) Perform weighted fusion on the data sets reconstructed each time based on the mean square error M and the mean absolute percentage error MAPE of the SAE feature reconstruction model:

[0158] 7.3.1) Calculate the collaborative error of the 5 SAE feature reconstruction models corresponding to each category of data sets:

[0159]

[0160] where E j represents the collaborative error of the j-th SAE feature reconstruction model in the i-th category of data sets based on the j-th mean square error M and the j-th mean absolute percentage error MAPE, and j is an integer in [1, 5];

[0161] 7.3.2) Calculate the weight of each SAE feature reconstruction model in each category of data sets according to the collaborative error E j :

[0162]

[0163] where w j is the weight of the i-th category of data sets in the j-th SAE feature reconstruction model, and n is the number of SAE feature reconstruction models corresponding to each category of data sets;

[0164] 7.3.3) According to the weight w j of each SAE feature reconstruction model in each category of data sets, fuse the feature reconstruction results of each category of data to obtain the new feature F a of each category of data sets:

[0165]

[0166] Among them, x is the reconstruction result of the data features of the current category, and a is an integer in [1, 5];

[0167] 7.4) Combine the new features F of the data sets of each category a That is, splice the corresponding columns of the features to obtain the fusion result F.

[0168] The above description is only a specific example of the present invention and does not constitute any limitation to the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these corrections and changes based on the idea of the present invention are still within the scope of protection of the claims of the present invention.

Claims

1. A high-dimensional feature reconstruction and fusion method based on SMT quality big data, characterized in that It includes the following steps: (1) Obtain the text data and structured data in the SMT production line quality big data, preprocess the text data and structured data respectively, and obtain the preprocessed SMT production line text data set and SMT production line structured data set; (2) Extract the features of the preprocessed SMT production line pre-text data set and SMT production line structured data set respectively to obtain a text feature set and a data feature set; (3) Merge and deduplicate the text feature set and the data feature set, and screen the original structured data set according to the merged and deduplicated feature set. Divide the screened data set into four categories: process, quality, production, equipment, and other data sets except these four categories; (4) Construct 5 integrated stacked autoencoder SAE feature reconstruction models with different activation functions, including an encoder and a decoder: Establish an SAE feature reconstruction model with Tanh as the activation function; Establish an SAE feature reconstruction model with Sigmoid as the activation function; Establish an SAE feature reconstruction model with Relu as the activation function; Establish an SAE feature reconstruction model with Softmax as the activation function; Establish an SAE feature reconstruction model with ReLU6 as the activation function; (5) Iteratively train the SAE feature reconstruction model to obtain the feature reconstruction result: (5a) Input the process, quality, production, equipment, and other data sets except these four categories into 5 different SAE feature reconstruction models, and output the predicted values of the quality indicators obtained for each type of data in each SAE feature reconstruction model; (5b) Set the loss function MSE for each SAE feature reconstruction model, and calculate the loss value J of the SAE feature reconstruction model based on the predicted and actual values of the quality metrics k ; (5c) Calculate the loss value J through the backpropagation method k of the loss gradient Then adopt the stochastic gradient descent method, and through the loss gradient update the weights w of the encoder and decoder in the SAE feature reconstruction model k until J k < 0.1, then stop training, and take the output result of the last iteration as the feature reconstruction result of each class of data; (6) Fuse the feature reconstruction results of each type of data: (6a) Take the loss value J obtained in step (5b) k as the mean square error M of each SAE feature reconstruction model in each category of data, and obtain the mean absolute percentage error MAPE of each SAE feature reconstruction model in each category according to the predicted value of the quality index corresponding to each category of data in step (5a); (6b) Based on the MSE and MAPE of the SAE feature reconstruction model, perform weighted fusion on the data set reconstructed each time to obtain the fusion result: (6b1) Calculate the collaborative error of each type of SAE feature reconstruction model; Among them, E j represents the collaborative error of the j-th SAE feature reconstruction model in each type of dataset based on the j-th mean square error M and the j-th mean absolute percentage error MAPE, where j is an integer in [1, 5]; (6b2) Calculate the weight of each type of SAE feature reconstruction model in each type of data set according to the collaborative error: where w j is the weight of each type of dataset in the j-th SAE feature reconstruction model, and n is the number of SAE feature reconstruction models corresponding to each type of dataset; (6b3) According to the weights of the reconstruction models of each type of SAE features in each type of dataset, fuse the feature reconstruction results of each type of data to obtain the new feature F of each type of dataset a : Where x is the feature reconstruction result of the data of the current category, and a is an integer in [1,5]; (6c) Combine the new features F of the datasets of each category a to obtain the fusion result F.

2. The method according to claim 1, wherein In the above (1), the preprocessing of the text data is realized as follows: Delete the irrelevant data in each text data set, and delete the serial numbers and spaces irrelevant to the key information; Fill in the missing numerical values in the text data in combination with business knowledge; Cut the text according to the nearest full stop and exclamation mark. Perform ontology construction, entity annotation, and entity relationship annotation on the text data after deletion and filling in sequence to obtain the preprocessed text data set.

3. The method according to claim 1, characterized in that, In the above (1), the preprocessing of the structured data is to first fill in the missing values in the structured data, then detect and delete the outliers in the filled data, and finally perform Z-score standardization on the filled and deleted data to obtain the preprocessed structured data set.

4. The method according to claim 1, characterized in that, In the above (2), the extraction of the features of the preprocessed SMT production line pre-text data set to obtain the text feature set is realized as follows: (2a) Construct a BERT-Bi-LSTM-CRF named entity recognition model composed of a BERT embedding layer, a Bi-LSTM layer, and a CRF layer connected in series; (2b) Train the BERT-Bi-LSTM-CRF named entity recognition model to obtain the named recognition result: Take the preprocessed SMT production line pre-text dataset as the input. According to the stochastic gradient descent method, based on the loss value of all feature vector sequences, backpropagate to adjust the number of neurons in the Bi-LSTM layer until the loss value is less than or equal to 0.

1. Take the output result of the last iteration as the named recognition result; (2c) Take the named recognition result as the input. According to the stochastic gradient descent method, based on the loss value of all feature vector relationship sequences, adjust the learning rate and dropout rate until the loss value is less than or equal to 0.

1. Take the output result of the last iteration as the knowledge extracted from the text data; (2d) Form the SMT production line quality knowledge graph according to the knowledge extracted from the text data; (2e) Input the SMT production line quality knowledge graph into the existing two-layer graph convolutional neural network GCN model, and use the stochastic gradient descent method to train it. Adjust the learning rate based on the loss value of each word vector sequence until the loss value is less than or equal to 0.

1. Take the model weights of the last iteration as the final weights of the model: (2f) Sort the weights obtained in (2e), and take the features with weights greater than or equal to 0.5 as the features extracted from the text dataset to obtain the text feature set.

5. The method according to claim 1, characterized in that, The (2) extracts the features of the preprocessed SMT production line structured dataset to obtain the data feature set, which is implemented as follows: (2g) Set the initialization parameters of the XGBoost feature extraction model; (2h) Input the features in the preprocessed SMT production line structured dataset into the XGBoost feature extraction model, and output the predicted values of the dataset quality indicators respectively; (2i) Use the importance formula in the XGBoost model to calculate the importance of the influencing factors of each feature in the dataset; (2j) Combine the features with the maximum influencing factor importance greater than 100 into the data feature set.

6. The method according to claim 1, characterized in that The text feature set extracted in (2) includes: Printing distance, demolding speed, demolding distance, squeegee length, squeegee pressure, belt speed, nitrogen concentration, number of manual cleaning times, and different temperatures of ten temperature zones, namely temperature of zone one, temperature of zone two, temperature of zone three, temperature of zone four, temperature of zone five, temperature of zone six, temperature of zone seven, temperature of zone eight, temperature of zone nine, temperature of zone ten.

7. The method according to claim 1, characterized in that, The data feature set extracted in (2) includes: Squeegee pressure, squeegee speed, printing height compensation, workbench separation speed, automatic cleaning count, cleaning speed, workbench separation distance, cleaning supply time, and squeegee separation distance.

8. The method according to claim 1, characterized in that, The encoder and decoder in the integrated stacked autoencoder SAE feature reconstruction model formed in (4) have the following structure parameters: For the encoder, it has two layers. The number of hidden neurons in the first input layer is set according to the number of features after merging and processing, and the number of hidden neurons in the second hidden layer is less than the number of hidden neurons in the first layer; The decoder has two layers. The number of hidden neurons in the first hidden layer is the same as that in the second hidden layer of the encoder, and the number of hidden neurons in the second output layer is less than that in the first hidden layer.

9. The method according to claim 1, characterized in that, In (5b), the loss function MSE of each SAE feature reconstruction model is set, and the formula is expressed as follows: where N is the number of each type of dataset, represents the predicted value of the quality index of each type of dataset, y i represents the actual value of the quality index of each type of dataset.

10. The method according to claim 1, wherein: The loss value J is calculated by the backpropagation method in (5c). k The loss gradient is given by the following formula: where w k is the weight of the k-th encoder and decoder; The random descent gradient method is adopted in (5c), and the weights w of the encoder and decoder in the SAE feature reconstruction model are updated through the loss gradient as follows: k ​ where w k ' represents w k the updated result, α represents the learning rate, and α ∈ [0, 1].

11. The method according to claim 1, wherein In (6a), the mean absolute percentage error MAPE of each SAE feature reconstruction model for each class is calculated, and the formula is as follows: Among them, N is the number of each type of dataset, represents the predicted value of the quality index of each type of dataset, y i represents the actual value of the quality index of each type of dataset.

Citation Information

Patent Citations

  • Data processing method and platform suitable for system security operation and maintenance

    CN115793590A

  • Radio frequency hopping signal identification method and system, terminal and medium

    CN115795302A

  • Knowledge graph construction method based on SMT quality big data analysis

    CN115098703A

  • Transfer learning-based method for improved VGG16 network pig identity recognition

    WO2022252272A1