Soft measurement method in industrial process and electronic equipment
By updating the historical gradient data of the output layer weight in the soft measurement data model, the problem of the traditional model decreasing prediction accuracy under dynamic changing conditions is solved, and higher accuracy and adaptability are achieved, and soft measurements in industrial processes are suitable.
Patent Information
- Application Number
- CN202510451686.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional offline training models cannot be adjusted in time when facing dynamically changing industrial processes, resulting in a decrease in prediction accuracy.
In the soft measurement data model, the historical gradient data of the output layer weight is determined based on the acquired data samples and prediction data, and the output layer weight is updated based on these historical gradient data to achieve online learning and adaptability of the model.
It improves the accuracy and prediction accuracy of soft measurement methods, can quickly adapt to data drift in industrial processes, reduce calculation overhead, and improve model stability and adaptability.
Smart Images

Figure CN120372206A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial measurement, and more particularly, to a soft measurement method and an electronic device in an industrial process. Background Art
[0002] In many industrial processes, data often exhibits time-varying and non-steady characteristics, resulting in performance degradation of traditional offline models in practical applications. Most traditional soft measurement methods rely on static data sets, which lack adaptability for dynamically changing operating conditions. When the distribution of key parameters in an industrial process changes (i.e., data drift), the offline-trained model often fails to adjust in a timely manner, leading to a decrease in prediction accuracy. Summary of the Invention
[0003] Embodiments of this application provide a soft measurement method and an electronic device in an industrial process to at least solve the technical problem of low prediction accuracy of offline-trained models.
[0004] According to the first aspect of the embodiments of this application, a soft measurement method in an industrial process is provided, including: during the process of soft measurement according to a soft measurement data model, determining historical gradient data of the output layer weights of the soft measurement data model based on the acquired data samples and prediction data;
[0005] Updating the output layer weights of the soft measurement data model according to the historical gradient data;
[0006] Performing soft measurement according to the updated soft measurement data model.
[0007] According to the second aspect of the embodiments of this application, an electronic device is provided, including a memory and a processor, where the memory is used to store computer instructions or computer programs; the processor is used to call and execute the computer instructions or computer programs to implement the soft measurement method described in the first aspect.
[0008] By adopting this embodiment, the accuracy and prediction accuracy of soft measurement in an industrial process can be improved. Brief Description of the Drawings
[0009] Figure 1 is a flowchart of a soft measurement method in an industrial process provided by an embodiment of this application;
[0010] Figure 2 is a flowchart of a method for determining historical gradient data of the output layer weights of a soft measurement data model provided by an embodiment of this application;
[0011] Figure 3 is a flowchart of a method for updating the output layer weights of a soft measurement data model provided by an embodiment of this application;
[0012] Figure 4 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0013] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0014] It should be understood that the "multiple" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; the "and / or" herein is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms such as "first" and "second" do not limit the quantity and execution order, and the terms such as "first" and "second" do not necessarily limit to be different
[0015] In addition, the terms "include" and "have" and any of their variants are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0016] First, the terms related to the embodiments of the present application are introduced.
[0017] Online learning: The online learning (Online Learning) model can continuously adapt to new data by updating the learning algorithm in real time. This is crucial for data flow modeling in industrial processes, especially for non-steady-state and multi-variable coupled systems.
[0018] Soft Sensor / Soft Measurement is a technology that uses easily measurable auxiliary variables (such as temperature, pressure, etc.) to indirectly estimate the dominant variables that are difficult to measure directly (such as component concentration, reaction rate, etc.). Its essence is to replace the "hardware direct measurement" of traditional instruments with "software calculation" to solve the problem that key parameters in industrial processes cannot be monitored in real time.
[0019] The above introduces the terms involved in the embodiments of the present application. Next, the application scenarios of soft measurement in the field of sewage treatment are exemplarily introduced.
[0020] In the field of sewage treatment, soft sensing technology is an important tool that uses neural network models (or soft sensing data models) to estimate key water quality parameters that are difficult to measure directly or are costly to measure. It mainly includes data collection, data preprocessing, neural network model design, model training and tuning, and model application.
[0021] In the data collection stage, we first need to collect various easily measurable input variables (such as flow, pH value, dissolved oxygen concentration, etc.) and target output variables (such as chemical oxygen demand COD, 5-day biochemical oxygen demand BOD5, total nitrogen TN, total phosphorus TP, etc.) in the sewage treatment process. These data usually come from online sensors and laboratory analysis.
[0022] In the data preprocessing stage, steps such as missing value processing, outlier detection and processing, data standardization or normalization are included to ensure the quality of the input data.
[0023] During the neural network model design phase, algorithm selection and architecture design are carried out. For example, based on the nature of the problem and the characteristics of the data, a suitable machine learning algorithm is selected, and the network structure is determined, such as the number of hidden layers and the number of neurons in each layer.
[0024] Afterwards, model training, model evaluation and verification, model adjustment and optimization can be performed based on the collected data. Finally, a neural network model that predicts the target output variable based on the input variables collected in real time can be obtained, which can be applied to soft measurement in the sewage treatment process.
[0025] The above is an illustrative introduction to the application scenarios of the embodiments of the present application. In this application scenario or similar scenarios, there are some traditional soft measurement methods that rely on static data sets, which lack adaptability to dynamically changing working conditions. When the distribution of key parameters in the industrial process changes (i.e., data drift), the offline trained model often cannot be adjusted in time, resulting in a decrease in prediction accuracy.
[0026] Based on this, the embodiment of the present application provides a soft measurement method in an industrial process, such as Figure 1As shown, the method includes the following processing procedures.
[0027] 100: During the soft sensing based on the soft sensing data model, based on the acquired data samples and predicted data, determine the historical gradient data of the output layer weights of the soft sensing data model.
[0028] 102: Update the output layer weights of the soft sensing data model according to the historical gradient data.
[0029] 104: Perform soft sensing according to the updated soft sensing data model.
[0030] By using the method provided in this embodiment, the historical gradient data of the output layer weights is updated using data samples and predicted data, and then the output layer weights of the soft sensing data model are updated according to the historical gradient data of the output layer weights, which can effectively improve the accuracy of the soft sensing method and the prediction accuracy.
[0031] In addition to being applicable to the sewage treatment scenario mentioned above, the embodiments of the present invention and the embodiments below, unless otherwise specified, are not limited to the sewage treatment scenario, and can also be applied to many industrial processes such as industrial process control, energy management, environmental monitoring, etc., and can all achieve the technical effect of improving the accuracy of soft sensing and the prediction accuracy.
[0032] Optionally, in an implementation manner of this embodiment, in process 100, the data samples include measurement data in the industrial process and actual data corresponding to the measurement data, and the predicted data is predicted by the soft sensing data model based on the measurement data. Exemplarily, taking the application to sewage treatment as an example, the measurement data can be easily measurable variables such as flow rate, pH value, concentration, etc. (for example, variables that can be easily detected in real time and conveniently by equipment), and the predicted data can be water quality indicators such as BOD5 (biochemical oxygen demand in 5 days). In other application scenarios, the measurement data can be variables such as temperature and pressure.
[0033] Exemplarily, for the measurement data in the industrial process, in order to ensure the stability and accuracy of the neural network model modeling, the measurement data is first normalized so that each feature has a zero mean and unit variance, avoiding the influence of the same feature scale difference on the model training. An example of the normalization process is as follows:
[0034]
[0035] where μ i , σ i are the mean and standard deviation of the i-th feature respectively, and x t,i represents the observed value of the i-th feature (or variable) at the t-th time point.
[0036] Optionally, in one implementation of this embodiment, for the soft sensor data model, given the input data x t , the calculation of the hidden layer h t = f(Wx t + b), where f(·) is the activation function, and ReLU, Sigmoid or Tanh is selected. The selection of the hidden layer activation function can be adjusted according to the characteristics of industrial data (such as industries like chemical engineering and metallurgy). W ∈ R h×d is the weight matrix of the hidden layer, and b ∈ R h is the bias.
[0037] Among them, h t = (h t,1 , h t,2 , …, h t,l ) is the output of the hidden layer, l is the size of the hidden layer, and W is the hidden layer weight. The calculation of the output layer is obtained by linearly combining the output of the hidden layer, and the calculation formula is:
[0038]
[0039] Among them, β = (β1, β2, …, β h ) is the weight of the output layer to be trained, h t is the output of the hidden layer, is the predicted value of the model (i.e., the predicted data).
[0040] In this implementation, the soft sensor data model is a feedforward neural network model and has fixed (not updated during the soft sensor process) hidden layer weights and biases. In this way, during the training of the soft sensor data model and the process of soft sensing, only the weights of the output layer can be trained and updated, thus greatly reducing the computational cost and being able to quickly adapt to new data streams. Exemplarily, the hidden layer weight W can be randomly initialized, and the elements of W can come from a uniform distribution or a Gaussian distribution.
[0041] Optionally, in one implementation of this embodiment, in process 100, as Figure 2 shown, the following method can be used to determine the historical gradient data of the output layer weights of the soft sensor data model.
[0042] 200: Calculate the value of the loss function of the soft sensor data model based on the predicted data and the actual data.
[0043] For example, for each sample (x t , y t ), y t represents the actual value corresponding to the predicted value. First, calculate the gradient of the loss function of the soft sensor data model. Taking the mean square error (MSE) as an example, the loss function is:
[0044]
[0045] 202: Calculate the current gradient of each weight in the output layer weights of the soft sensor data model based on the value of the loss function.
[0046] For example, the gradient of the loss function with respect to the output layer weight β is:
[0047]
[0048] 204: For any one of the weights, update the historical gradient data of any one of the weights according to the current gradient of any one of the weights.
[0049] Exemplarily, the historical gradient data includes a gradient cumulative value and a squared gradient cumulative value. After determining the current gradient, the gradient cumulative value and the squared gradient cumulative value can be determined.
[0050] Optionally, in one implementation manner of this embodiment, in process 102, updating the output layer weights of the soft sensor data model according to the historical gradient data includes: updating each weight according to the historical gradient data of each weight in the output layer weights of the soft sensor data model. Specifically, as Figure 3 shown, it includes:
[0051] 300: For any one of the weights, determine the value of the first regularization function of any one of the weights according to the first historical gradient data of any one of the weights, and the first regularization function is used to control the sparsity of the output layer weights.
[0052] 302: Determine the value of the second regularization function of any one of the weights according to the second historical gradient data of any one of the weights, and the second regularization function is used to control the stability of the output layer weights.
[0053] 304: Calculate the updated value of any one of the weights according to the value of the first regularization function and the value of the second regularization function. For example, use the ratio of the value of the first regularization function to the value of the second regularization function as the updated value of any one of the weights.
[0054] In this implementation manner, the first historical gradient data and the second historical gradient data are different data used to reflect the historical information of the gradient.
[0055] Adopting this implementation manner, the change information of the gradient is reflected through different historical gradient data, and the sparsity and stability of the output layer weights are balanced through the first regularization function and the second regularization function, which is beneficial to reducing the complexity of the model, improving the calculation efficiency, and improving the stability.
[0056] Using this implementation method for online updating of the soft sensor data model, different from the traditional Stochastic Gradient Descent (SGD) method, this implementation method improves the stability of the update by introducing a regularization term and historical gradient information. Especially when facing data drift, it can quickly adapt to the new data distribution.
[0057] Optionally, in this implementation method, the first historical data is the gradient cumulative value of any weight, and the first regularization function is the L1 regularization function for the gradient cumulative value; the second historical data is the gradient square cumulative value of any weight, and the second regularization function is the L2 regularization function for the gradient square cumulative value.
[0058] For example, the first regularization function f1 is expressed as: when λ1 - |z i | < 0, f1 = -sign(z i )·(λ1 - |z i |), when λ1 - |z i | ≥ 0, f1 = 0; the second regularization function is expressed as: The update value of any weight is expressed as:
[0059] where z i represents the gradient cumulative value of the i-th weight, λ1 is the L1 regularization coefficient, sign(z i ) represents the sign of z i , α represents the learning rate, n i represents the gradient square cumulative value of the i-th weight, λ2 represents the L2 regularization coefficient, and β i represents the update value of the i-th weight.
[0060] where z i ←z i +g i -σ i β i ; n i ←n i +g i 2 ; g i is the gradient of β i .
[0061]
[0062] In other embodiments, it can also be adopted: z i ←z i +g i ; n i ←n i +g i 2 ; g i is the gradient of βi Gradient.
[0063] Sparsifying the weights through L1 regularization means that some weights will be forced to zero, effectively reducing the complexity of the model and improving the computational efficiency. By adjusting the value of λ1, the degree of sparsification can be controlled. Exemplarily, λ1 typically takes values between 0 and 1. A smaller λ1 value (such as 0.01 to 0.1) can produce a sparser model, while a larger λ1 value (such as 0.1 to 1) will further increase the sparsity of the model. λ2 typically takes values between 0 and 0.1. A smaller λ2 value (such as 0.001 to 0.01) can provide weak regularization, while a larger λ2 value (such as 0.01 to 0.1) will enhance the regularization effect.
[0064] In an embodiment of the present application, the soft sensor data model is a feedforward neural network model. In the soft sensor method, without changing the weights and biases of the hidden layer of the soft sensor data model, the weights of the output layer of the soft sensor data model are updated according to the historical gradient data. Exemplarily, the soft sensor data model can be an RVFL (Random Vector Functional Link Network) model, which randomly initializes the weights of the hidden layer and only trains the weights of the output layer, thereby reducing the computational complexity and improving the generalization ability of the model.
[0065] In addition, by combining the method for updating the weights of the output layer in the soft sensor method provided in each embodiment of the present application, the problems of parameter update and data drift faced by the traditional RVFL model during online learning can be solved. For example, the Figure 3 embodiment shown in the present application provides an adaptive, stable and efficient update mechanism for online learning, which can effectively cope with the data drift problem. Compared with the traditional SGD (Stochastic Gradient Descent) and other online learning methods, the advantages are at least one of the following:
[0066] (1) Sparsity processing: It can achieve L1 regularization during the update process, forcing some weights of the output layer to tend to zero, thereby improving the sparsity of the model. This sparsity helps to improve the generalization ability of the model and avoid overfitting. For the RVFL model, this sparse update can effectively reduce the computational burden during the training process and help the model maintain high efficiency when data flows in.
[0067] (2) Ability to cope with data drift: By accumulating gradient information and regularization techniques, it can efficiently cope with concept drift and feature drift. When the data stream in the industrial process continuously changes, it can help the RVFL model quickly adjust the weights of the output layer, enabling it to maintain good prediction performance under the new data distribution.
[0068] (3) Efficient computation and real-time update: Different from traditional batch learning or offline training methods, the embodiments of this application can be updated when each sample arrives, reducing computational latency and being particularly suitable for industrial applications that require real-time prediction. Compared with traditional online learning methods (such as SGD), by introducing a regularization term, the model becomes more robust and can handle more drastic industrial processes.
[0069] In other words, although some existing technologies have proposed variants of online RVFL, most of them have the following limitations: (1) Low learning efficiency: Many existing online RVFL models rely on simple online gradient descent algorithms, lacking historical accumulation and adaptive adjustment of gradient information. Such an approach has poor performance in dealing with data drift and consumes a large amount of computational resources. (2) Unable to handle sparsity: Most existing online RVFL models fail to effectively achieve sparsification, resulting in an overly complex model, large computational overhead, and difficulty in ensuring the efficiency of the model. (3) Poor adaptability: In the face of complex industrial data streams, many online RVFL models fail to quickly adapt to the dynamic changes of data, especially unable to quickly adjust to new data distributions, resulting in a decrease in model prediction accuracy. By adopting the embodiments provided in this application (for example, Figure 3 the embodiments shown), these problems can be well optimized.
[0070] To more clearly highlight the beneficial effects of the embodiments Figure 3 shown in this application, the following will compare and explain the scheme of updating RVFL in the embodiments Figure 3 shown in this application with the existing scheme of updating the RVFL model using least squares (ALS). There is an existing technology that uses alternating least squares (ALS) to alternately update the weights of the linear and non-linear parts of the RVFL model until the set error or iteration count requirement is met, thereby obtaining the final network model. In contrast, updating the RVFL model using the embodiments Figure 3 shown in this application has the following advantages:
[0071] (1) Strong real-time performance: The online update mechanism of FTRL can respond immediately to the data stream, avoiding the problem of complex calculations in each iteration of ALS. (2) High computational efficiency: The sparse update and adaptive learning rate of FTRL reduce computational overhead and are suitable for industrial applications with large-scale data streams. (3) Strong adaptability: FTRL can effectively handle data drift and ensure the stability of the model in a dynamic industrial environment. (4) Wide generalization: It is applicable to soft sensing in industrial processes such as chemical engineering and metallurgy and has better industrial application prospects.
[0072] Specifically, the embodiments of this application and the ALS update method have the following differences:
[0073] A. Model Update Mechanism:
[0074] ALS method: Alternately update the output weights of the linear and non - linear parts of the model. In each iteration, the parameters of each part are optimized by the least - squares method. Generally, this method fixes the parameters of one part first and then updates the other part in each step, so as to gradually approach the optimal solution. This method depends on the alternating iteration of the model parameters and may lead to a long calculation time, especially when dealing with large - scale data.
[0075] Embodiment of this application: During the online learning process, the output layer weights (i.e., the linear part) are updated in real - time, and through adaptive learning rate and sparsification update, the model can adjust the weights according to real - time data. The embodiment of this application is faster than ALS because it can update the weights immediately after each data input, rather than relying on the alternating iteration method. In addition, the sparsification feature of the embodiment of this application makes the update process more efficient, especially when dealing with large - scale and sparse data streams.
[0076] B. Computational Efficiency and Real - time Performance:
[0077] The ALS method usually needs to fix a part of the parameters and update the other part in each iteration. The calculation involves matrix inversion and error iteration, which will increase the calculation time, especially when there are many model parameters. Although this method can ensure the convergence of the model, it may lead to low computational efficiency in the online learning environment of dynamic data streams.
[0078] The embodiment of this application has higher computational efficiency because it adjusts the weights step by step. Each time, only the output layer parameters need to be updated, without complex operations such as matrix inversion. Therefore, it is suitable for online learning, especially for quickly responding in real - time data streams and industrial process modeling.
[0079] C. Sparsification and Regularization:
[0080] The ALS method mainly optimizes the output weights by the alternating least - squares method and may not be able to effectively control overfitting, especially on large - scale data sets, which may lead to too high model complexity.
[0081] The embodiment of this application introduces L1 regularization, which enables the weight update process to force some parameters to approach zero, thus sparsifying the model. This not only reduces the computational overhead but also improves the generalization ability of the model.
[0082] 2. Advantage Analysis:
[0083] A. Adaptability: The embodiments of the present application optimize the performance of RVFL in the face of dynamically changing data streams and data drift. Through the accumulation of historical gradients and an adaptive learning rate, the model parameters can be adjusted in real time to adapt to the continuous changes in industrial processes. In contrast, the ALS algorithm requires multiple iterations to gradually update the parameters, resulting in a slower response speed when dealing with dynamic changes.
[0084] B. Stability and Efficiency: The online learning mechanism of the embodiments of the present application allows the model to be updated quickly at each step of data stream input, ensuring that the model can stably handle the changes in data during industrial processes. The ALS method may be limited by convergence and computational efficiency during the iterative process. In cases where a large number of iterations are required, it consumes a large amount of computing resources and has poor real-time performance.
[0085] C. Physical Interpretability and Applicability: The ALS method makes the model more physically interpretable by alternately updating the linear and non-linear parts, which helps to understand the application of the model in actual production. However, this method has low computational efficiency and is difficult to adapt to large-scale data streams. The method of the embodiments of the present application does not lose physical interpretability during model updates, but through sparsification and online learning techniques, it provides higher computational efficiency and adapts to large-scale industrial production processes in industries such as chemical engineering and metallurgy.
[0086] D. Model Generalizability: The embodiments of the present application have stronger generalizability in multiple industrial scenarios, especially for online learning and large-scale data stream scenarios. Its high efficiency enables the model to be widely applied to various dynamic non-steady-state systems. The ALS method, due to its dependence on alternating updates and iterative calculations, is mainly applicable to scenarios that require offline processing and model optimization, and it is difficult to be widely promoted in industrial applications of real-time data streams.
[0087] In Figure 3 the illustrated embodiment, in order to further improve the balance between the sparsity and stability of the output weights and achieve better model generalization ability, the L1 regularization coefficient during the calculation of each weight in the output layer weights is dynamically adjusted according to historical gradient information. Specifically, the L1 regularization coefficient (λ1) of the first regularization function used for the next update of any weight can be adjusted according to the gradient accumulation value of any weight. The following is a detailed description.
[0088] In this embodiment, adjusting the L1 regularization coefficient of the first regularization function used for the next update of any weight according to the gradient accumulation value of any weight includes: If the gradient accumulation value z corresponding to a certain parameter iMeeting the first set condition (for example, being higher than a set value, which can be the average of the gradient cumulative values of all parameters, or being high enough to a set situation, such as being high enough to belong to the top 10% of the values arranged from large to small) indicates that this parameter has a greater impact on the objective function and may require a stronger sparsity constraint. At this time, when calculating β i+1 a corresponding L1 regularization coefficient of this parameter can be increased. For example, a set value or a set amplitude is increased, and the set value or set amplitude is related to the degree to which the gradient cumulative value is higher than the set condition. For example, if z i belongs to the gradient cumulative values in the top 10%-5% (including 5%) of the highest values, then the first set value or the first set amplitude is increased; if z i belongs to the gradient cumulative values in the top 5%-2% (excluding 5%) of the highest values, then the second set value or the second set amplitude is increased; if z i belongs to the gradient cumulative values in the top 2% (excluding 2%) of the highest values, then the third set value or the third set amplitude is increased. Among them, the first set value, the second set value, and the third set value increase in sequence, and the first set amplitude, the second set amplitude, and the third set amplitude are percentages that increase in sequence. For the specific values of the set value and the set amplitude, this embodiment does not make specific limitations.
[0089] Alternatively, according to the gradient cumulative value of any weight, adjust the L1 regularization coefficient of the first regularization function used when updating any weight next time, including: if the gradient square cumulative value of any weight meets the first set condition, then determine the L1 regularization coefficient of the first regularization function used when updating any weight next time according to the L1 regularization coefficient update formula. Among them, the L1 regularization coefficient update formula is used to represent the corresponding relationship between the L1 regularization coefficient when updating any weight next time and the change amount of the gradient cumulative value after updating any weight in the current time.
[0090] Exemplarily, the L1 regularization coefficient update formula is expressed as: λ 1,t+1,i =k1·Δz t,i +m1; Δz t,i =z t,i -z t-1,i . Among them, k1 and m1 are adjustment factors, which can be optimized according to the actual situation and set as fixed values in actual applications. This embodiment does not limit their specific values. z t,i represents the gradient cumulative value obtained after the current update of the i-th output layer weight, z t-1,i represents the gradient cumulative value obtained after the previous update of the i-th output layer weight, and λ 1,t+1,i represents the L1 regularization coefficient of the i-th output layer weight at the (t + 1)-th update.
[0091] Optionally, in a practical application of this embodiment, the method for updating the L1 regularization coefficient provided in this embodiment can be applied after the first set number of data samples. That is, for the first set number of data samples, the L1 regularization coefficient and the L2 regularization coefficient can adopt set values without update, and only for the subsequent data samples are they updated, so as to further maintain the stability of the prediction accuracy of the RVFL model by adjusting the regularization coefficients and cope with the possible data drift problem after a large number of data samples. The specific value of the first set number is not limited in the embodiments of this application.
[0092] In other embodiments of this application, the soft sensor data model can use an Extreme Learning Machine (ELM) model, an Echo State Network (ESN) model, a Sparse Convolutional Network (SCN) model, etc. to replace the RVFL model, or use a deep learning network (for example, a convolutional neural network or a recurrent neural network) model to replace the RVFL model.
[0093] Figure 4 is a schematic diagram of an electronic device according to an embodiment of this application. As Figure 4 shown, the electronic device includes a processor 10, at least one communication bus 20, a user interface 30, at least one external communication interface 40, and a memory 50. Among them, the communication bus 20 is configured to realize the connection and communication between these components. Among them, the user interface 30 may include a display screen, and the external communication interface 40 may include a standard wired interface and a wireless interface. Among them, computer instructions or computer programs are stored in the memory 50. Among them, the processor 10 is used to execute the computer instructions or computer programs stored in the memory 50 to implement the method provided in the method embodiment of this application.
[0094] The description of the above electronic device is similar to the description of the above method embodiment and has similar beneficial effects to the method embodiment. For the technical details not disclosed in the electronic device of this application, please refer to the description of the method embodiment of this application for understanding.
[0095] The serial numbers or the order of introduction of the embodiments of this application are only for description and do not represent the superiority or inferiority of the embodiments.
[0096] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0097] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0098] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0099] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. For example, a computer program product includes one or more computer instructions, which when executed implement following the regularization leader algorithm to thus implement Figure 3The method provided by the illustrated embodiment. When computer instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by the computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid state disk (SSD)), etc. It should be noted that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, a non-transitory storage medium.
[0100] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the scene data of the current frame in the three-dimensional virtual scene, the device information of the client, and the scene interaction information involved in the embodiments of the present application are all obtained under full authorization.
[0101] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A soft measurement method in an industrial process, characterized in that The method includes: During the soft sensing based on the soft sensing data model, determining historical gradient data of the output layer weights of the soft sensing data model based on the acquired data samples and prediction data; Updating the output layer weights of the soft sensing data model according to the historical gradient data; Performing soft sensing according to the updated soft sensing data model.
2. The soft sensing method according to claim 1, wherein: The data samples include measurement data in an industrial process and actual data corresponding to the measurement data; The prediction data is predicted by the soft sensing data model based on the measurement data.
3. The soft measurement method according to claim 2, characterized in that, The determining historical gradient data of the output layer weights of the soft sensing data model based on the acquired data samples and prediction data includes: Calculating the value of the loss function of the soft sensing data model according to the prediction data and the actual data; Calculating the current gradient of each weight in the output layer weights of the soft sensing data model based on the value of the loss function; For any one of the weights, updating the historical gradient data of the any one of the weights according to the current gradient of the any one of the weights.
4. The soft measurement method according to claim 1, characterized in that The updating the output layer weights of the soft sensing data model according to the historical gradient data includes: Updating each weight according to the historical gradient data of each weight in the output layer weights of the soft sensing data model.
5. The soft sensing method according to claim 4, characterized in that, The updating each weight according to the historical gradient data of each weight in the output layer weights of the soft sensing data model includes: For any one of the weights, determining the value of the first regularization function of the any one of the weights according to the first historical gradient data of the any one of the weights, where the first regularization function is used to control the sparsity of the output layer weights; Determining the value of the second regularization function of the any one of the weights according to the second historical gradient data of the any one of the weights, where the second regularization function is used to control the stability of the output layer weights; Calculating the update value of the any one of the weights according to the value of the first regularization function and the value of the second regularization function.
6. The soft sensing method according to claim 5, wherein: The first historical data is the cumulative value of the gradients of the any one of the weights, and the first regularization function is the L1 regularization function with respect to the cumulative value of the gradients; The second historical data is the cumulative value of the squared gradients of the any one of the weights, and the second regularization function is the L2 regularization function with respect to the cumulative value of the squared gradients.
7. The soft sensing method according to claim 6, wherein: The first regularization function f1 is expressed as: f1 = -sign(z i )·(λ1 - |z i |), λ1 - |z i | < 0, f1 = 0, λ1 - |z i | ≥ 0; The second regularization function is expressed as: The updated value of any of the weights is expressed as: Among them, the z i represents the gradient cumulative value of the i-th weight, λ1 is the L1 regularization coefficient, sign(z i ) represents the sign of z i , α represents the learning rate, n i represents the square cumulative value of the i-th weight gradient, λ2 represents the L2 regularization coefficient, β i represents the update value of the i-th weight.
8. The soft measurement method according to claim 6 or 7, characterized in that The method further includes: Adjusting the L1 regularization coefficient of the first regularization function used for the next update of the any one of the weights according to the cumulative value of the gradients of the any one of the weights.
9. The soft sensing method according to claim 1, wherein: The soft sensing data model is a feedforward neural network model; The updating the output layer weights of the soft sensing data model according to the historical gradient data includes: updating the output layer weights of the soft sensing data model without changing the hidden layer weights and biases of the soft sensing data model.
10. An electronic device, characterized in that, Comprising a memory and a processor, wherein, the memory is used for storing computer instructions or computer programs; the processor is used for calling and executing the computer instructions or computer programs to implement the soft measurement method according to any one of claims 1-9.