A Large-Scale Data Mining Method, System and Storage Medium for Manufacturing Processes
Through the neural network algorithm and support vector machine algorithm with gradient accumulation compensation, the stability and efficiency problems of traditional models in multidimensional data processing are solved, faster convergence and higher robustness are achieved, and equipment fault diagnosis and production efficiency are improved.
Patent Information
- Application Number
- CN202510287308.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-12
AI Technical Summary
When traditional machine learning models process multidimensional data, dynamically changing data and noisy data, they are unable to effectively respond to challenges in industrial production, resulting in insufficient generalization capabilities of models, affecting equipment maintenance and production efficiency, and neglecting the problem of sample imbalance affecting the accuracy of fault diagnosis.
The neural network algorithm using gradient accumulation compensation is used to optimize parameter updates through gradient accumulation and dynamic compensation, combined with sample importance adjustment and local gradient offset calculation, dynamically adjust the learning rate and activation function, and classification is used by the support vector machine algorithm, dynamically balance training data categories, and incrementally update the model.
The convergence process of complex process data is accelerated, the model's ability to identify equipment status is improved, the fault diagnosis ability is enhanced, the model's adaptability and robustness is improved, the stable training of high-dimensional and nonlinear data is ensured, and the equipment maintenance and production efficiency is optimized.
Smart Images

Figure CN119806095B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent manufacturing, and in particular relates to a large-scale data mining method, system and storage medium for manufacturing processes. Background Art
[0002] With the increasing intelligence of industrial equipment and complexity of production processes, traditional machine learning models are often unable to effectively cope with the challenges encountered in actual production when processing multidimensional data, dynamically changing data, and noisy data, resulting in insufficient generalization of the model, which in turn affects equipment maintenance, fault diagnosis, and production efficiency. Currently, data in many industrial applications usually involves a large number of feature points and has problems such as high dimensionality, redundancy, and noise.
[0003] The Chinese invention patent with the prior art publication number CN119089141A proposes a method and system for optimizing water treatment based on machine learning. The method first collects process variable data such as water quality, flow, liquid level, and pressure through sensors arranged at different parts of the water treatment equipment; the collected process variable data are stored according to type and structure storage to realize data partitioning and indexing; then the stored data is cleaned and standardized, a CNN-LSTM algorithm model is constructed, and the data is stream processed through the CNN-LSTM algorithm model; finally, the results predicted by the model are compared with the preset rules, and the control strategy is confirmed according to the comparison results. According to the confirmed control strategy, the water treatment equipment in each process link is controlled to achieve the overall water treatment effect. The method realizes the accurate prediction, real-time monitoring and scheduling of water treatment process variables, realizes the optimization goals, energy efficiency, and automatic operation and unattended operation of the water treatment system, and significantly improves the level of intelligence, processing efficiency and response speed.
[0004] However, the training process may encounter difficulties in convergence due to abnormal changes in gradients, and may even fail to effectively capture key information in the data. Especially in the processing of complex and multi-dimensional industrial data, existing methods have low stability and efficiency when processing high-dimensional data.
[0005] At the same time, the problem of sample imbalance is ignored, resulting in minority samples being often ignored or not fully learned, which may affect the accuracy of fault diagnosis and further affect production efficiency and equipment operation and maintenance level. Summary of the invention
[0006] The present invention proposes a large-scale data mining method for manufacturing processes, which realizes parameter update optimization by means of gradient accumulation compensation, thereby overcoming the gradient explosion and gradient vanishing problems that are prone to occur in traditional neural networks in large-scale process data training.
[0007] In order to solve the above problems, the technical solution provided by the present invention is:
[0008] In a first aspect, a large-scale data mining method for manufacturing processes includes data collection, where the data is obtained from manufacturing processes;
[0009] Extract features from the data, where a neural network algorithm based on gradient accumulation compensation is used as a feature extraction model; the neural network based on gradient accumulation compensation is obtained through training;
[0010] Classify the data according to the features, where a classification model is established using a machine learning algorithm to identify the data;
[0011] Store and manage the classified data;
[0012] Visualize the data.
[0013] Furthermore, the training method of the neural network algorithm based on gradient accumulation compensation includes,
[0014] Initialization of the parameters of the neural network, where the parameters of the neural network are initialized by means of a random normal distribution, and the initial value of gradient accumulation is set to zero; to avoid unstable phenomena such as premature convergence or oscillation of the model in the initial state, during the backpropagation process of each training, a gradient accumulation strategy is adopted to add and accumulate the current gradient and the historical gradient, and a dynamic compensation factor is used for smooth update; the calculation method of the gradient accumulation strategy is expressed as:
[0015]
[0016] In the formula, is the gradient accumulation of the m-th iteration, is the gradient accumulation of the (m - 1)-th iteration, m is a positive integer, β pe is the gradient accumulation decay factor, is the gradient of the loss function of the neural network with respect to the parameter θ p , is the m-th training sample; is the label of the m-th training sample; L(·, ·) is the loss function of the neural network; is the local gradient offset for the m-th iteration; θ p is the neural network parameter, and the neural network parameter includes the weight and bias parameters of the neural network; through gradient accumulation compensation, the model can better capture these small changes and improve the ability to identify the device state. By accumulating the past data gradients, the model can be more sensitive to the initial signals of abnormalities in the production process;
[0017] Local gradient offset calculation, where the calculation of the local gradient offset performs weighted processing on important samples through sample weighting coefficients;
[0018] Update the neural network parameter gradient using the loss function of the neural network;
[0019] During the forward calculation of the neural network, dynamically adjust the activation function, and combine the activation output with the scale of the current cumulative gradient information and the input sample features;
[0020] Dynamically regulate the learning rate according to the gradient change during the training process;
[0021] Globally update the neural network parameters;
[0022] Repeat the above steps iteratively. If the continuous change amplitude of the loss function is less than the given threshold, stop training and complete the neural network training based on gradient accumulation compensation.
[0023] Furthermore, the calculation method for dynamically adjusting the activation function is expressed as:
[0024]
[0025] In the formula, is the activation value of the nth sample in the neural network, n is a positive integer, Sig(·) is the Sigmoid activation function, E p is the number of neurons in the neural network, α pse is the first forward propagation hyperparameter, β pse is the second forward propagation hyperparameter, is the cumulative gradient of the mth iteration, is the nth training sample, is the norm of the global features in the training dataset, || || is the L2 norm, is the eth weight of the neural network, e is a positive integer, is the eth bias of the neural network.
[0026] Furthermore, dynamically regulate the adaptive learning rate based on the adaptive learning rate adjustment strategy for batch gradient norm statistics; the calculation method of the adaptive learning rate adjustment strategy based on batch gradient norm statistics is expressed as:
[0027]
[0028] In the formula, is the learning rate of the tth iteration of the neural network, η0 is the initial learning rate of the neural network, N p is the number of samples input to the neural network in the current batch, r is a positive integer, β pve is the learning rate adjustment factor, τp is the threshold for learning rate adjustment, is the gradient accumulation at the t-th iteration.
[0029] Furthermore, the update method of the loss function of the neural network for the parameter gradient is expressed as:
[0030] In the formula, ← is the parameter update operation, γ p is the gradient update adjustment factor, τ p is the preset gradient threshold, δ pa is the hyperparameter for controlling the non-linear scaling of the gradient.
[0031] Furthermore, the classification model structure is based on the support vector machine algorithm.
[0032] Furthermore, it is characterized in that the data obtained in the manufacturing process includes production equipment sensor data, production process monitoring system data, equipment maintenance system data, quality inspection equipment data, and manual record data.
[0033] In a second aspect, a large-scale data mining method for manufacturing processes is developed into a large-scale data mining system for manufacturing processes, including
[0034] a data acquisition module, which is used to acquire data in the manufacturing process;
[0035] a data processing module, which is used to extract features from the data;
[0036] a data classification module, which is used to classify the data according to the features;
[0037] a data storage and management module, which is used to store and manage the classified data;
[0038] a data visualization and analysis module, which outputs the visualized processing of the data.
[0039] In a third aspect, a computer-readable storage medium storing instructions, in which a computer program or instructions are stored, and when the computer program or instructions are executed by an image processing device, the above-mentioned large-scale data mining method for manufacturing processes is implemented.
[0040] In a fourth aspect, a computer program product is characterized in that the computer program product includes: computer program code, and when the computer program code runs, it causes a processor to execute the above-mentioned large-scale data mining method for manufacturing processes.
[0041] Advantages of the present invention:
[0042] 1. The present invention uses a neural network algorithm based on gradient accumulation compensation as a feature extraction model, and realizes the update and optimization of parameters through gradient accumulation compensation, thereby overcoming the problems of gradient explosion and gradient disappearance that are prone to occur in the training of large-scale process data by traditional neural networks, and further accelerating the convergence process of complex process data, showing stronger adaptability and robustness.
[0043] 2. Through gradient accumulation compensation, the model can better capture these tiny changes and improve the ability to identify the equipment status. By accumulating the past data gradients, the model can be more sensitive to the initial signals of abnormalities occurring in the production process.
[0044] 3. By adjusting the current gradient according to the sample importance, the calculation of the local gradient offset is weighted for important samples through the sample weighting coefficient, so as to better capture the key features related to the equipment maintenance system. The data in the equipment maintenance system often includes some rare fault information. The samples of these information are few but extremely important. By dynamically adjusting the sample weights, the model can improve the recognition rate of these rare events, and further optimize the fault diagnosis ability of the equipment.
[0045] 4. Through the implicit gradient truncation mechanism, excessive or unstable gradients will be truncated, thus ensuring the stable training of the neural network in high-dimensional and non-linear data. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a comparison chart of training losses of different optimization methods;
[0047] Figure 2 It is a comparison chart of gradient stabilities;
[0048] Figure 3 It is a comparison chart of the accuracies of comparing mainstream models such as ResNet and LSTM on the industrial sensor dataset;
[0049] Figure 4 It is a comparison chart of classification accuracies under different noises. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0052] The present invention proposes a large-scale data mining method for manufacturing processes, which mainly includes the following modules:
[0053] (1) Data acquisition module
[0054] The function of the data acquisition module is to obtain relevant raw data from the manufacturing process. The data sources in the present invention are mainly industrial data generated during the manufacturing process. Various machine devices, sensors and all links of the production line in the manufacturing process will generate a large amount of data, which contains information about equipment status, environmental parameters, production efficiency, etc. The data acquisition sources include:
[0055] 1) Production equipment sensors: During operation, the equipment records physical quantities such as temperature, pressure, rotation speed, and vibration through sensors;
[0056] 2) Production process monitoring system: Real-time monitoring data records key nodes in the production process, including operating parameters, raw material usage, production progress, etc.;
[0057] 3) Equipment maintenance system: Records data such as equipment maintenance history, fault logs, and repair times;
[0058] 4) Quality inspection equipment: Inspection data generated by automated inspection equipment, such as product dimensions, surface quality, etc.;
[0059] 5) Manually recorded data: Data recorded through manual inspections, which supplements data that cannot be covered by sensors or when abnormalities occur.
[0060] The data acquisition method is through the combination of industrial control systems and Internet of Things devices, which automatically collects raw data from various sensors and control systems in real time and transmits it to the data storage platform via wireless or wired means.
[0061] The collected data is saved in a distributed database in JSON format. Each data record contains multi-dimensional information such as device ID, timestamp, operating parameters, status information, etc., and is structured and stored in a predefined standard format.
[0062] In order to enable the data mining model to achieve supervised training, the collected data is labeled manually. The labeling categories include:
[0063] 1) Category 1: Abnormal state, equipment failure, process deviation, environmental mutation, etc.
[0064] 2) Category 2: Normal state, the process is in normal state.
[0065] (2) Feature Engineering Module
[0066] The function of the feature engineering module is to use neural network technology to extract effective features from data and improve the performance of machine learning models.
[0067] The present invention adopts a neural network algorithm based on gradient accumulation compensation as a feature extraction model, and realizes parameter update optimization by means of gradient accumulation compensation, thereby overcoming the gradient explosion and gradient vanishing problems that are prone to occur in traditional neural networks in large-scale process data training, thereby accelerating the convergence process of complex process data and showing stronger adaptability and robustness.
[0068] Specifically, the training process of the neural network algorithm based on gradient accumulation compensation is as follows:
[0069] 1. Initialize the parameters of the neural network by random normal distribution, and set the initial value of the gradient accumulation to zero to avoid the model from converging prematurely or oscillating instability in the initial state. For example, in the process data such as temperature and pressure collected by the production equipment sensor, due to the large fluctuation of the initial data, the traditional model may be overfitted or unstable due to the fluctuation of the process data. The initialization strategy of the present invention can effectively avoid this problem. The initialization method is expressed as:
[0070]
[0071]
[0072] In the formula, is the e-th weight of the neural network, e is a positive integer, It means the mean is 0 and the variance is The normal distribution of is the e-th bias of the neural network, is the initialization variance of the neural network, and ~ indicates that it obeys a specific distribution. Preferably, Set to 0.1.
[0073] Furthermore, during the backpropagation process of each training, a gradient accumulation strategy is adopted. The current gradient is added and accumulated with the historical gradient, and smoothed updates are performed through a dynamic compensation factor to reduce the instability caused by the drastic fluctuations of the gradient of a single sample. For example, for the time series data in the production process monitoring system, the monitoring system may experience instantaneous data fluctuations. Through gradient accumulation compensation, the model can better capture these tiny changes, improve the ability to identify the device state, and make the model more sensitive to the initial signals of abnormalities in the production process by accumulating the past data gradients. The calculation method is expressed as:
[0074]
[0075] In the formula, is the gradient accumulation of the m-th iteration, is the gradient accumulation of the (m - 1)-th iteration, m is a positive integer, and β pe is the gradient accumulation decay factor, is the gradient of the loss function of the neural network with respect to the parameter θ p , is the m-th training sample; is the label of the m-th training sample; L(·, ·) is the loss function of the neural network; is the local gradient offset for the m-th iteration; θ p is the neural network parameter, including the weights and bias parameters of the neural network. Preferably, β pe is set to 0.95.
[0076] It should be noted that the conventional technology does not adopt the gradient accumulation strategy. When dealing with large-scale process data, problems such as gradient disappearance or gradient explosion often occur, which will cause the model to converge slowly or even fail to converge during the training process. In the process data feature extraction task, such as the prediction model of the equipment maintenance system, due to the existence of a large number of non-linear features in the process data, it is difficult for traditional methods to effectively solve these problems, resulting in low feature extraction efficiency and even the inability to effectively capture the key state of the equipment.
[0077] Furthermore, the current gradient is adjusted by sample importance. The calculation of the local gradient offset performs weighted processing on important samples through the sample weighting coefficient, so as to better capture the key features related to the equipment maintenance system. The data in the equipment maintenance system often includes some rare fault information. The samples of these information are few but extremely important. By dynamically adjusting the sample weights, the model can improve the recognition rate of these rare events, and then optimize the fault diagnosis ability of the equipment. For example, some extreme fault modes may occur less frequently in historical data, and traditional models often ignore these data. After adopting the weighting strategy, the model can pay more attention to these "rare but important" fault modes during the learning process. The calculation method of the local gradient offset is expressed as:
[0078]
[0079] In the formula, α p is the sample weighting coefficient, || || is the L2 norm, is the norm of the current gradient, λ pev is the regularization gain factor; is the feature mean of the m-th training sample; σ pev is the standard deviation hyperparameter of the sample. Preferably, α p is set to 0.2, λ pev is set to 0.1, and σ pev is set to 0.2.
[0080] Furthermore, an implicit gradient truncation mechanism is adopted to judge and truncate excessive or unstable gradients, ensure the stability of high-dimensional process data training, and prevent gradient explosion or excessive shrinkage in complex models. The process data generated by quality inspection equipment is often very high-dimensional and noisy. Traditional neural networks may be unstable during training due to the interference of some abnormal process data points. Through the implicit gradient truncation mechanism, excessive or unstable gradients will be truncated, thus ensuring the stable training of the neural network in high-dimensional and non-linear data. For example, when modeling the quality inspection data of chemical products, some inspection equipment may cause data anomalies due to improper operation or external environmental influence. Through this mechanism, the interference of these abnormal process data on model training can be effectively avoided. The update method of the parameter gradient by the loss function of the neural network is expressed as:
[0081]
[0082] In the formula, ← is the parameter update operation, γ p is the gradient update adjustment factor, τ p is the preset gradient threshold, δ pa is the hyperparameter that controls the non-linear scaling of the gradient. Preferably, γ p is set to 0.5, δpa Set to 2, τ p Set to 0.1.
[0083] 2. During the feedforward calculation of the neural network, dynamically adjust the activation function so that the activation output combines the current gradient accumulation information and the scale of the input sample features, realizing a more flexible mapping of high-dimensional and non-linear process data. Production process monitoring data usually has strong non-linear characteristics. Especially under the influence of multi-variable cross-effects, traditional methods may not be able to fully capture this complexity. By combining the current gradient information and input features, the model can more flexibly adapt to complex data and optimize the mapping effect of the data. For example, in the production monitoring of equipment, the change of equipment status is often affected by the comprehensive influence of multiple factors. By dynamically adjusting the activation function, the complex relationship between these multiple factors can be better revealed. The calculation method is expressed as:
[0084]
[0085] In the formula, is the activation value of the nth sample in the neural network, n is a positive integer, Sig(·) is the Sigmoid activation function, E p is the number of neurons in the neural network, α pse is the first forward propagation hyperparameter, β pse is the second forward propagation hyperparameter, is the gradient accumulation of the mth iteration, is the nth training sample, is the norm of the global features in the training dataset, || || is the L2 norm. Preferably, α pse is set to 0.1, β pse is set to 0.3.
[0086] It should be noted that when dealing with high-dimensional and non-linear data, traditional neural network models usually adopt fixed activation functions, which easily lead to insufficient mapping ability of the model when facing complex process data and unable to effectively capture the non-linear features in the data. Especially in quality inspection equipment, there may be complex non-linear relationships in the data, and fixed activation functions cannot flexibly adapt to high-dimensional and non-linear data, resulting in the model being unable to effectively map the complex features in the production process and restricting the performance of the model.
[0087] 3. An adaptive learning rate adjustment strategy based on batch gradient norm statistics is adopted to dynamically regulate the learning rate according to the gradient changes during the training process, ensuring a faster and more stable convergence rate in complex scenarios. Due to the real-time nature and variability of the monitoring data in the production process, traditional fixed learning rate methods may perform poorly in the face of complex scenarios. By dynamically adjusting the learning rate, the convergence speed can be accelerated while maintaining the stability of training. For the equipment maintenance system, dynamic adjustment during the training process can ensure that when facing sudden abnormal process data of equipment, the model can quickly adjust the learning strategy. The calculation method is expressed as:
[0088]
[0089] In the formula, is the learning rate of the neural network at the t-th iteration, η0 is the initial learning rate of the neural network, N p is the number of samples input to the neural network in the current batch, r is a positive integer, β pve is the learning rate adjustment factor, τ p is the threshold for learning rate adjustment, is the gradient accumulation at the t-th iteration. Preferably, η0 is set to 0.001, β pve is set to 2, τ p is set to 0.95.
[0090] 4. Globally update the network parameters to globally converge to a better solution. The calculation method is expressed as:
[0091]
[0092] In the formula, is the parameter of the neural network at the t-th iteration, is the parameter of the neural network at the (t + 1)-th iteration, is the learning rate of the neural network at the t-th iteration, is the gradient accumulation at the t-th iteration.
[0093] 5. Repeat the above steps iteratively. Determine whether the training can be terminated according to the stability of the training error. If the continuous change amplitude of the loss function is less than the given threshold, it is considered that the model has converged and the training is stopped. The calculation method is expressed as:
[0094]
[0095] In the formula, ∈ pcg is the preset allowable minimum error change amount, is the (m - 1)-th training sample, is the label of the (m - 1)-th training sample; is the label of the m-th training sample; is the label of the m-th training sample; L(·, ·) is the loss function of the neural network; preferably, ∈ pcg is set to 0.001.
[0096] Furthermore, to verify the effectiveness of this method, through experimental data analysis, as follows Figure 1 shown, the training loss comparison graph shows that this method has a faster convergence speed compared to optimization methods such as Adam, SGD, and RMSprop, and the slope of the convergence curve of this method is larger and it enters the stable stage earlier.
[0097] In addition, as follows Figure 2 shown, the gradient stability analysis graph shows that the gradient norm of this method is always stable within a reasonable range, and the convergence process is more stable.
[0098] In addition, as follows Figure 3 shown, when comparing with mainstream models such as ResNet and LSTM on the industrial sensor dataset, this method has the highest accuracy (94.3%) and the smallest standard deviation (±1.2%), and has higher robustness.
[0099] In addition, as follows Figure 4 shown, through noise robustness analysis, it shows that as the input noise increases, the accuracy of this method decreases more slowly, and still maintains an accuracy of more than 75% under the noise intensity of 0.5, which is better than traditional methods.
[0100] (3) Machine learning modeling module
[0101] The role of the machine learning modeling module is to perform classification tasks based on data, use machine learning algorithms to establish a classification model, and help identify different categories and patterns in the data.
[0102] (4) Data storage and management module
[0103] The role of the data storage and management module is to efficiently store and manage a large amount of data generated by the system, ensuring the security, accessibility, and scalability of the data.
[0104] The functions of the data storage and management module include: efficient data storage, supporting distributed storage and fast reading of large-scale data; data backup and recovery, ensuring that data will not be lost in case of accidents and can be quickly restored; data permission management, ensuring that the access and operation permissions of data comply with security standards and preventing illegal access; data indexing and retrieval, improving data query efficiency through indexing technology to meet the needs of fast data retrieval.
[0105] The machine learning classifier model structure adopted by the present invention is based on the support vector machine algorithm, and the training process is as follows:
[0106] 1. During the training process of the support vector machine, it is necessary to first process the redundant features in the training data. The present invention adopts a redundancy elimination mechanism based on information entropy. By calculating the correlation and redundancy degree between features, the feature space of the data set is dynamically adjusted. Process data usually contains high-dimensional and interrelated features, such as sensor data of equipment temperature, pressure, vibration, etc. If directly input into the support vector machine model, it will bring redundancy, thus affecting the performance of the model. The redundancy elimination loss function can automatically identify and delete redundant features, retain the features that have the greatest impact on the classification result, reduce the computational complexity, and improve the classification accuracy. The calculation method of the redundancy elimination loss function is expressed as:
[0107]
[0108] In the formula, is the redundancy elimination loss function; I(X c,i , X c,j ) is the mutual information between feature X c,i and X c,j , measuring the correlation between the two; λ ij is the redundancy coefficient of feature X c,i and X c,j , adaptively adjusted according to its importance; X c,i and X c,j respectively represent the i-th and j-th features input into the support vector machine.
[0109] Furthermore, the calculation method of the redundancy coefficient is expressed as:
[0110]
[0111] In the formula, α ij is the redundancy measurement between feature X c,i and X c,j ; α ik is the redundancy measurement between feature X c,i and X c,k ; X c,k represents the k-th feature input into the support vector machine.
[0112] Furthermore, the calculation method of the redundancy measurement is expressed as:
[0113]
[0114] σ ij =max(‖X c,i ‖, ‖X c,j ‖)
[0115] In the formula, ‖‖ represents the L2 norm; max(,) is the maximum value function.
[0116] 2. To address the problem of class imbalance, the present invention adopts a dynamic balance strategy, enabling the class ratio of the training data to be adaptively adjusted as the training process progresses. Suppose there are N A and N B samples of class A and class B respectively in the current dataset, and class B is significantly fewer than class A. Then the class weights W class,A and W class,B are adjusted as follows:
[0117]
[0118]
[0119] In the formula, W class,A and W class,B are the weights of class A and class B respectively; N max is the number of samples in the class with more samples; N A and N B are the number of samples of class A and class B respectively.
[0120] Furthermore, during the training process, the samples with larger class weights are input into the support vector machine earlier in the batch. The dynamic balance strategy can effectively prevent the model from being biased towards the class with more samples and improve the classification accuracy of the minority class.
[0121] 3. During the training process of the support vector machine, selecting which samples as support vectors is crucial for the model performance. The traditional support vector machine adopts globally optimal support vectors. The present invention dynamically selects the most representative support vectors through a regularization term, and the optimization objective function is:
[0122]
[0123] In the formula, is the loss function of the support vector machine; N u is the number of samples in the training set; y u,i is the label of the i-th sample; x u is the feature vector of the i-th sample; w u is the weight of the support vector machine model; b u is the bias of the support vector machine model; λ u is the regularization term coefficient of the support vector machine. Preferably, λ u is set to 0.2.
[0124] 4. To enable the model to meet the requirements of incremental learning, the present invention adopts an incremental update strategy. This strategy can update the support vectors in real time as new data is added. After each new data is adopted, the model adjusts the weights through the following incremental update formula:
[0125]
[0126] wherein, is the weight of the updated support vector machine model; is the weight of the support vector machine model before update; η u is the learning rate of the support vector machine. Preferably, η u is set to 0.01.
[0127] 5. After all the above optimization steps are completed, the final support vector machine model is trained and optimized through the loss functions of all steps, and the final objective function is:
[0128]
[0129] wherein, is the total loss function of the support vector machine. By minimizing this loss function, the optimal support vector machine classifier model is trained, which can efficiently perform data classification tasks.
[0130] (5) Data Visualization and Analysis Module
[0131] The function of the data visualization and analysis module is to convert complex data results and model outputs into easy-to-understand graphical information to help users make decision-making analysis.
[0132] The functions of the data visualization and analysis module include: real-time monitoring and display, visualizing real-time production data, model prediction results, and process status through dashboards, etc.; data trend analysis, displaying the change trends of data in the form of charts, heat maps, etc. to help users identify potential problems; prediction result display, presenting the prediction results of the machine learning model (such as equipment failure prediction, production capacity prediction, etc.) in an intuitive way for production personnel to make timely responses.
[0133] Compared with the prior art, the present invention adopts a neural network algorithm based on gradient accumulation compensation, which can effectively avoid the problems of gradient explosion and gradient disappearance encountered by traditional neural networks in large-scale process data training. Through gradient accumulation and dynamic compensation, the model can stably update parameters and accelerate the convergence process of complex process data, and has stronger adaptability and robustness.
[0134] A calculation method for sample importance adjustment and local gradient offset is proposed to weight rare and important samples to improve the model's recognition ability for these samples. Especially in the fault diagnosis of the equipment maintenance system, for those rare but extremely important fault modes, the model can pay more attention and learn.
[0135] Adopting an incremental learning strategy, it can update the model in real time according to new data, enabling the model to flexibly adjust parameters when facing continuously changing production data, thus avoiding the limitations of the fixed training mode.
[0136] Adopting a redundancy elimination mechanism based on information entropy, by calculating the correlation between features, it dynamically adjusts the feature space, removes redundant features, improves the computational efficiency and classification accuracy of data, enabling the model to retain the most informative features and avoiding the performance degradation caused by redundant features.
[0137] Those of ordinary skill in the art will realize that the embodiments described herein are to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.
Claims
1. A large-scale data mining method for manufacturing processes, characterized in that, Including: Data acquisition, where the data is data obtained in the manufacturing process; Extracting features from the data, where a neural network algorithm based on gradient accumulation compensation is used as the feature extraction model; Among them, the training method of the neural network algorithm based on gradient accumulation compensation includes: Initializing the parameters of the neural network, where the parameters of the neural network are initialized by means of random normal distribution, and the initial value of gradient accumulation is set to zero; the neural network parameters include the weights and bias parameters of the neural network; In the backpropagation process of each training, a gradient accumulation strategy is adopted to add and accumulate the current gradient and the historical gradient, and a dynamic compensation factor is used for smooth update; the calculation method of the gradient accumulation strategy is expressed as: In the formula, is the gradient accumulation of the m-th iteration, is the gradient accumulation of the (m - 1)-th iteration, m is a positive integer, and β pe is the gradient accumulation decay factor, is the gradient of the loss function of the neural network with respect to the parameter θ p ; is the m-th training sample; is the label of the m-th training sample; L(·, ·) is the loss function of the neural network; is the local gradient offset for the m-th iteration; θ p is the neural network parameter; where the calculation method of the local gradient offset is expressed as: In the formula, α p is the sample weighting coefficient, |||| is the L2 norm, is the norm of the current gradient, and λ pev is the regularization gain factor; is the feature mean of the m-th training sample; σ pev is the standard deviation hyperparameter of the sample; Calculating the local gradient offset, and the calculation of the local gradient offset performs weighted processing on important samples through sample weighting coefficients; Updating the neural network parameter gradient using the loss function of the neural network; During the feedforward calculation of the neural network, the activation function is dynamically adjusted so that the activation output combines the current cumulative gradient information and the scale information of the input sample features; Dynamically regulating the learning rate according to the gradient change situation during the training process; Globally updating the neural network parameters; Repeatedly iterate the above steps, and stop training if the continuous change amplitude of the loss function is less than a given threshold; Classifying the data according to the features, where a classification model is established using a machine learning algorithm to identify the data; among them, the classification model structure is based on the support vector machine algorithm, and the training of the support vector machine algorithm includes: Processing redundant features in the training data, where based on the redundancy elimination mechanism of information entropy, by calculating the correlation and redundancy between features, the feature space of the data set is dynamically adjusted; Automatically identifying and deleting redundant features using the redundancy elimination loss function; Storing and managing the classified data; Performing visualization processing on the data.
2. The large-scale data mining method for manufacturing processes according to claim 1, characterized in that The calculation method for dynamically adjusting the activation function is expressed as: In the formula, is the activation value of the nth sample in the neural network, n is a positive integer, Sig(·) is the Sigmoid activation function, E p is the number of neurons in the neural network, a pse is the first forward propagation hyperparameter, β pse is the second forward propagation hyperparameter, is the gradient accumulation of the mth iteration, is the nth training sample, is the norm of the global features in the training dataset, || || is the L2 norm, is the e-th weight of the neural network, e is a positive integer, is the e-th bias of the neural network.
3. A large-scale data mining method for manufacturing processes according to claim 1, characterized in that Dynamically regulating the adaptive learning rate based on the adaptive learning rate adjustment strategy for batch gradient norm statistics; the calculation method of the adaptive learning rate adjustment strategy for batch gradient norm statistics is expressed as: In the formula, is the learning rate of the t-th iteration of the neural network, η0 is the initial learning rate of the neural network, N p is the number of samples input to the neural network in the current batch, r is a positive integer, β pve is the learning rate adjustment factor, τ p is the threshold for learning rate adjustment, is the gradient accumulation of the t-th iteration.
4. A large-scale data mining method for manufacturing processes according to claim 1, characterized in that The update method of the neural network loss function for the parameter gradient is expressed as: where ← is the parameter update operation, γ p is the gradient update adjustment factor, τ p is the preset gradient threshold, δ pa is the hyperparameter that controls the non-linear scaling of the gradient.
5. A large-scale data mining method for manufacturing processes according to claim 1, characterized in that The data obtained in the manufacturing process includes production equipment sensor data, production process monitoring system data, equipment maintenance system data, quality inspection equipment data, and manual record data.
6. A computer-readable storage medium for storing instructions, characterized in that, The storage medium stores a computer program or instruction, and when the computer program or instruction is executed by an image processing device, the method described in any one of claims 1-5 is implemented.
7. A computer program product, characterized in that, The computer program product includes computer program code, and when the computer program code runs, it causes the processor to execute any one of the methods in claims 1-5.
Citation Information
Patent Citations
Optimized water treatment method and system based on machine learning
CN119089141A
Network data mining method and device, equipment and storage medium
CN119311942A