Method and apparatus, device, and medium for determining latent variables of Gaussian mixture model
By setting the sample weight matrix and responsiveness matrix in the Gaussian mixed model and updating the hidden variables, the problem of sample imbalance in industrial predictions is solved, the early warning effect and judgment ability are improved, and the introduction of noise is avoided.
Patent Information
- Application Number
- CN202210460579.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-28
AI Technical Summary
In industrial predictive maintenance, Gaussian mixed models have deviated early warning effects due to imbalance in sample size. The prior art such as SMOTE is not effective in high-dimensional space and is prone to introduce noise.
By setting the initial value of the hidden variable, determine the sample weight matrix according to the number of samples, calculate the response matrix and update the hidden variable until the convergence conditions are met, consider the sample weight differences, avoid artificial data insertion, and increase the weight of a few types of samples.
It improves the early warning effect of the Gaussian hybrid model, solves the problem of sample imbalance, avoids the introduction of noise, is suitable for working conditions with a small number of samples in industrial scenarios, and improves the ability to judge normal working conditions and potential risks.
Smart Images

Figure CN114817844B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial prediction, and particularly relates to a method and device, equipment, and medium for determining latent variables of a Gaussian mixture model. Background Art
[0002] In the application practice of industrial predictive maintenance, when using a Gaussian mixture model to fit the normal operation mode in the actual production scenario, in most cases, there will be an imbalance in the number of samples. This is because there are often different working conditions in production, and the running time of each working condition is not exactly the same. Therefore, some working conditions will generate more samples, while some working conditions will generate fewer samples, resulting in the problem of sample imbalance. For the working conditions with fewer samples, the early warning effect of the model will be biased.
[0003] Currently, the main solution to the above problem is the Synthetic Minority Oversampling Technique (SMOTE for short), that is, to solve the imbalance problem by expanding the minority class samples. Essentially, this technique increases the data volume by randomly inserting artificial data between the data points that already exist in the fewer number of samples. However, such a method has poor effects in high-dimensional spaces and will introduce noise in the case where the working conditions are relatively close. Therefore, the current solution cannot well solve the problem of sample imbalance in industrial scenarios. Summary of the Invention
[0004] The present invention provides a method and device, equipment, and medium for determining latent variables of a Gaussian mixture model, which can better solve the problem of sample imbalance in industrial scenarios.
[0005] In a first aspect, an embodiment of the present invention provides a method for determining latent variables of a Gaussian mixture model, including:
[0006] S1. Set an initial value of the latent variable based on the training sample set of the Gaussian mixture model;
[0007] S2. Determine a sample weight matrix corresponding to the training sample set according to the number of samples corresponding to each working condition in the training sample set; wherein, the sample weight matrix includes the sample weights of each training sample in the training sample set; in each working condition corresponding to the training sample set, the fewer the number of samples corresponding to a working condition, the greater the sample weights of each training sample corresponding to this working condition;
[0008] S3. Determine the response matrix corresponding to the current iteration process according to the sample weight matrix and the current value of the latent variable. The response matrix includes the responses of each training sample under each working condition to each Gaussian kernel of the Gaussian mixture model. The response of a training sample to a Gaussian kernel is the expected probability that the training sample belongs to the Gaussian kernel.
[0009] S4. Update the current value of the latent variable according to the response matrix corresponding to the current iteration process.
[0010] S5. Determine whether the current value of the updated latent variable satisfies the convergence condition.
[0011] If so, end this method and use the updated current value of the latent variable as the optimal solution.
[0012] Otherwise, return to S3 to perform the next iteration process.
[0013] Optionally, the determining the sample weight matrix corresponding to the training sample set according to the number of samples corresponding to each working condition in the training sample set includes:
[0014] Calculate the sample weight of the nth training sample in the training sample set using the first calculation formula, and the first calculation formula includes:
[0015]
[0016] In the formula, W n is the sample weight of the nth training sample, s in is the number of samples corresponding to the ith working condition to which the nth training sample belongs, s j is the number of samples corresponding to the jth working condition, and I is the number of each working condition corresponding to the training sample set.
[0017] Optionally, the latent variable includes the kernel weight of each Gaussian kernel and the kernel parameters of each Gaussian kernel, and the kernel parameters include the mean matrix and the covariance matrix.
[0018] Further, the determining the response matrix corresponding to the current iteration process according to the sample weight matrix and the current value of the latent variable includes:
[0019] Calculate the response of the nth training sample to the kth Gaussian kernel in the response matrix corresponding to the current iteration process using the second calculation formula, and the second calculation formula includes:
[0020]
[0021] In the formula, r(z nk) is the responsiveness of the nth training sample in the responsiveness matrix corresponding to this iteration process to the kth Gaussian kernel, W n is the sample weight of the nth training sample, π k is the current value of the kernel weight of the kth Gaussian kernel, x n is the nth training sample, μ k is the current mean matrix of the kth Gaussian kernel, ∑ k is the current covariance matrix of the kth Gaussian kernel, K is the number of Gaussian kernels, and N() is the Gaussian function.
[0022] Optionally, updating the current value of the latent variable according to the responsiveness matrix corresponding to this iteration process includes at least one of the following:
[0023] For each Gaussian kernel, calculate the updated kernel weight of the Gaussian kernel according to the sample weights of each training sample and the responsiveness of each training sample in the responsiveness matrix corresponding to this iteration process to the Gaussian kernel;
[0024] For each Gaussian kernel, calculate the updated mean matrix of the Gaussian kernel according to the responsiveness of each training sample in the responsiveness matrix corresponding to this iteration process to the Gaussian kernel and each training sample in the sample training set;
[0025] For each Gaussian kernel, calculate the updated covariance matrix of the Gaussian kernel according to each training sample in the sample training set, the updated mean matrix of the Gaussian kernel, and the responsiveness of each training sample in the responsiveness matrix corresponding to this iteration process to the Gaussian kernel.
[0026] Further, calculating the updated kernel weight of each Gaussian kernel according to the sample weights of each training sample and the responsiveness of each training sample in the responsiveness matrix corresponding to this iteration process to the Gaussian kernel includes:
[0027] Calculate the updated kernel weight of the kth Gaussian kernel using the third calculation formula, and the third calculation formula includes:
[0028]
[0029] where, π k is the updated kernel weight of the kth Gaussian kernel; r(z nk ) is the responsiveness of the nth training sample in the responsiveness matrix corresponding to this iteration process to the kth Gaussian kernel; N is the number of training samples in the training sample set; W n is the sample weight of the nth training sample.
[0030] Further, for each Gaussian kernel, according to the responsiveness of each training sample in the responsiveness matrix corresponding to the current iteration process to this Gaussian kernel and each training sample in the sample training set, calculate the updated mean matrix of this Gaussian kernel, including:
[0031] Use the fourth calculation formula to calculate the updated mean matrix of the k-th Gaussian kernel, and the fourth calculation formula includes:
[0032]
[0033] In the formula, μ k is the updated mean matrix of the k-th Gaussian kernel, x n is the n-th training sample, r(z nk ) is the responsiveness of the n-th training sample in the responsiveness matrix corresponding to the current iteration process to the k-th Gaussian kernel, and N is the number of training samples in the training sample set.
[0034] Further, for each Gaussian kernel, according to each training sample in the sample training set, the updated mean matrix of this Gaussian kernel, and the responsiveness of each training sample in the responsiveness matrix corresponding to the current iteration process to this Gaussian kernel, calculate the updated covariance matrix of this Gaussian kernel, including:
[0035] Use the fifth calculation formula to calculate the updated covariance matrix of the k-th Gaussian kernel, and the fifth calculation formula includes:
[0036]
[0037] In the formula, ∑ k is the updated covariance matrix of the k-th Gaussian kernel, N is the number of training samples in the training sample set, x n is the n-th training sample, r(z nk ) is the responsiveness of each training sample in the responsiveness matrix corresponding to the current iteration process to this Gaussian kernel, and μ k is the updated mean matrix of the k-th Gaussian kernel.
[0038] In the second aspect, an embodiment of the present invention provides a device for determining latent variables of a Gaussian mixture model, including:
[0039] An initialization module, configured to set an initial value of the latent variable based on the training sample set of the Gaussian mixture model;
[0040] A weight determination module, configured to determine a sample weight matrix corresponding to the training sample set according to the number of samples corresponding to each working condition in the training sample set; wherein, the sample weight matrix includes the sample weights of each training sample in the training sample set; among the various working conditions corresponding to the training sample set, the fewer the number of samples corresponding to a working condition, the greater the sample weights of the respective training samples corresponding to this working condition;
[0041] A responsiveness determination module, configured to determine a responsiveness matrix corresponding to the current iteration process according to the sample weight matrix and the current value of the latent variable; wherein, the responsiveness matrix includes the responsiveness of each training sample under each working condition to each Gaussian kernel of the Gaussian mixture model, and the responsiveness of a training sample to a Gaussian kernel is the expected probability that this training sample belongs to this Gaussian kernel;
[0042] An update module, configured to update the current value of the latent variable according to the responsiveness matrix corresponding to the current iteration process;
[0043] A convergence judgment module, configured to judge whether the current value of the updated latent variable satisfies the convergence condition; if so, end the method provided by this device, and use the current value of the updated latent variable as the optimal solution; otherwise, return to the responsiveness determination module to execute the next iteration process.
[0044] In a third aspect, an embodiment of the present invention provides a computing device, which includes: at least one memory and at least one processor;
[0045] The at least one memory is used to store machine-readable programs;
[0046] The at least one processor is configured to call the machine-readable program and execute the method provided in the first aspect.
[0047] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the processor is caused to execute the method provided in the first aspect.
[0048] The latent variable determination method, device, equipment, and medium of the Gaussian mixture model provided in the embodiment of the present invention first initialize the latent variables, then determine the sample weight matrix according to the number of samples of each working condition, and then determine the responsiveness matrix corresponding to this iterative process according to the sample weight matrix and the latent variables, and then use the responsiveness matrix of this iterative process to update the latent variables. If the updated latent variables meet the convergence conditions, the method is terminated and the current value of the updated latent variables is used as the optimal solution, that is, the current value of the latent variables at this time is the optimal result; otherwise, the next iterative process is entered until the optimal solution is obtained. In this scheme, the sample weights of each training sample are considered when calculating the responsiveness matrix, and the sample weights of each training sample under the condition with fewer samples are higher, that is, the difference between the number of samples corresponding to different working conditions is taken into account, and the training samples under the working condition with fewer samples have higher sample weights. In this way, after convergence, the kernel weights of each Gaussian kernel have a certain bias, and the standard solution results tend to give the minority class samples a higher weight of the local optimal result, so that the working condition with a small number of samples has a certain decision-making power in the subsequent fitting of the Gaussian mixture model, and will not be considered as an abnormal class due to the small number of samples, thereby improving the judgment of normal working conditions and potential risks in industrial scenarios, that is, the latent variables calculated by this scheme can improve the early warning effect of the Gaussian mixture model. Moreover, this scheme does not need to increase the number of samples corresponding to the working condition by inserting artificial data into each sample under the working condition with a small number of samples, and will not introduce noise when the working conditions are relatively close. It can be seen that this scheme can better solve the problem of sample imbalance in industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0050] Figure 1 It is a flowchart of a method for determining latent variables of a Gaussian mixture model provided by an embodiment of the present invention;
[0051] Figure 2 It is a structural block diagram of a latent variable determination device for a Gaussian mixture model provided by an embodiment of the present invention.
[0052] Description of reference numerals:
[0053] S1~S5 Step 100 Latent variable determination device for Gaussian mixture model 110 Initialization module 120 Weight determination module 130 Responsiveness determination module 140 Update module 150 Convergence judgment module DETAILED DESCRIPTION
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] In a first aspect, an embodiment of the present invention provides a method for determining latent variables of a Gaussian mixture model, and this method can be executed by any computing device. Refer to Figure 1 , and this method includes the following steps S1 to S4:
[0056] S1. Set an initial value of the latent variable based on the training sample set of the Gaussian mixture model;
[0057] Among them, the Gaussian mixture model, namely GMM, is an extension of a single Gaussian probability density function, and uses multiple Gaussian probability density functions (i.e., normal distribution curves) to accurately quantify the variable distribution. Therefore, the Gaussian mixture model is a statistical model that can decompose the variable distribution into several distributions based on Gaussian probability density functions. Each Gaussian probability density function is a Gaussian model, and a Gaussian model is a Gaussian kernel, that is, a Gaussian mixture model has multiple Gaussian kernels.
[0058] Among them, the training sample set includes training samples corresponding to multiple working conditions, and the number of samples corresponding to each working condition may be different. Some working conditions have more corresponding samples, while some working conditions have fewer corresponding samples. The so-called working condition is the normal operation mode under different scenarios. For example, daytime working conditions and nighttime working conditions. During the day, the working hours of the device are 10 hours, while at night the working hours of the device are 2 hours. It can be seen that there are more training samples under daytime working conditions, and fewer training samples at night. For another example, for a certain device, there are high-speed operation mode, medium-speed operation mode, and low-speed operation mode, that is, the device corresponds to three working conditions, and the number of operation data of the device under each working condition is different. For another example, for a production enterprise, there are multiple pump groups, and it can be decided which pump groups are in the on state according to needs. Different numbers of pump groups in the on state correspond to different working conditions. For another example, a device sensitive to temperature corresponds to four working conditions when operating in the four seasons of spring, summer, autumn, and winter; for another example, the production processing line has working conditions such as full load, half load, and 70% load.
[0059] Among them, the training samples are the operation data of the device in a specific scenario. For example, a training sample is a vector formed by multiple operation data (such as multiple parameters like rotational speed, device temperature, etc.) of a certain device at a certain time point. The device can be a key device with multiple sensors, so that the operation data of the device can be collected according to each sensor. Using the trained Gaussian mixture model can fit the working conditions of the device under various working conditions, and realize the prediction of the working conditions of the device in the future period. Since the Gaussian mixture model has high robustness, it can well fit the working conditions of the device under multiple working conditions. Before using the Gaussian mixture model to fit the operation conditions of the device, it is necessary to determine the latent variables of the Gaussian mixture model.
[0060] Among them, the so-called latent variables can include the kernel weights of each Gaussian kernel, the mean matrix and / or covariance matrix of each Gaussian kernel, etc. The mean matrix and covariance matrix are the kernel parameters of the Gaussian kernel. It can be seen that the latent variables of a Gaussian kernel are the related variables of this Gaussian kernel.
[0061] In specific implementation, assume that the training sample set is X. Take the covariance matrix of the training sample set X as the covariance matrix of each Gaussian kernel, and evenly distribute the kernel weights to each Gaussian kernel, that is, the initial values of the kernel weights of each Gaussian kernel are the same. Randomly select multiple training samples in the training sample set, and calculate the mean matrix of these multiple training samples as the mean matrix of a Gaussian kernel. In this way, the mean matrices of all Gaussian kernels can be obtained. It can be seen that the initial values of the latent variables of each Gaussian kernel can be obtained in this way. Of course, before determining the initial values of the latent variables, the samples in the training sample set can be normalized first to facilitate subsequent calculations.
[0062] Among them, the number of Gaussian kernels of the Gaussian mixture model can be determined and adjusted according to the complexity of the device and the fitting effect of the Gaussian mixture model. The number of Gaussian kernels is different from the number of working conditions and also different from the number of parameters to be fitted by the Gaussian mixture model. However, the number of Gaussian kernels is related to the number of working conditions. Generally speaking, the number of Gaussian kernels is greater than or equal to the number of working conditions. On this basis, it is further adjusted by the fitting effect of the Gaussian mixture model.
[0063] S2. Determine the sample weight matrix corresponding to the training sample set according to the number of samples corresponding to each working condition in the training sample set; among them, the sample weight matrix includes the sample weights of each training sample in the training sample set; in each working condition corresponding to the training sample set, the fewer the number of samples corresponding to a working condition, the greater the sample weights of each training sample corresponding to this working condition;
[0064] That is to say, in S2, the sample weights of each training sample are calculated, and the sample weights of each training sample form the above sample weight matrix. For example, if there are N training samples in the training sample set, a sample weight matrix with a dimension of N*1 can be obtained. Moreover, the sample weights of the training samples corresponding to different working conditions are different. If the number of samples included in the working condition to which a training sample belongs is smaller, then the sample weight of this training sample is larger.
[0065] In specific implementation, in S2, the first calculation formula can be used to calculate the sample weight of the nth training sample in the training sample set, and the first calculation formula includes:
[0066]
[0067] In the formula, W n is the sample weight of the nth training sample, s in is the number of samples corresponding to the ith working condition to which the nth training sample belongs, s j is the number of samples corresponding to the jth working condition, and I is the number of each working condition corresponding to the training sample set.
[0068] It can be seen that the sample weights of each training sample can be calculated through the above first calculation formula. In the sample weight matrix calculated through the above first calculation formula, the sample weights of each training sample under the same working condition are the same, and the sample weights of the training samples under different working conditions are different. However, if the number of samples corresponding to a working condition is smaller, the sample weights of each training sample under this working condition are larger. If the number of samples corresponding to a working condition is larger, then the sample weights of each training sample under this working condition are smaller. That is to say, it can be seen from the above first calculation formula that the sample weights of each training sample under each working condition are inversely proportional to the number of samples corresponding to this working condition.
[0069] S3. Determine the response matrix corresponding to the current iteration process according to the sample weight matrix and the current value of the latent variable; wherein, each training sample under each working condition in the response matrix respectively corresponds to the response degrees of each Gaussian kernel of the Gaussian mixture model, and the response degree of a training sample to a Gaussian kernel is the expected probability that this training sample belongs to this Gaussian kernel;
[0070] It can be understood that each training sample in the training sample will have K response degrees for K Gaussian kernels, that is, a training sample corresponds to K response degrees. In this way, the dimension of the response matrix is N*K, where N is the number of training samples in the training sample set, and K is the number of Gaussian kernels of the Gaussian mixture model. A training sample corresponds to a row of response degrees in the response matrix.
[0071] Among them, the responsiveness of a training sample to a Gaussian kernel is the expected probability that the training sample belongs to the Gaussian kernel. It can be seen that S3 is actually a process of calculating the expected value.
[0072] It can be understood that if this iteration process is the first iteration process, the current value of the latent variable is the initial value of the latent variable. If this iteration process is not the first iteration process, the current value of the latent variable is the value after updating the latent variable in the previous iteration process. In short, the current value of the latent variable is the latest value of the latent variable.
[0073] In specific implementation, in S3, the second calculation formula can be used to calculate the responsiveness of the nth training sample in the responsiveness matrix corresponding to this iteration process to the kth Gaussian kernel. The second calculation formula includes:
[0074]
[0075] In the formula, r(z nk ) is the responsiveness of the nth training sample in the responsiveness matrix corresponding to this iteration process to the kth Gaussian kernel, W n is the sample weight of the nth training sample, π k is the current value of the kernel weight of the kth Gaussian kernel, x n is the nth training sample, μ k is the current mean matrix of the kth Gaussian kernel, ∑ k is the current covariance matrix of the kth Gaussian kernel, K is the number of Gaussian kernels, and N() is the Gaussian function.
[0076] Similarly, μ j is the current mean matrix of the jth Gaussian kernel, ∑ j is the current covariance matrix of the jth Gaussian kernel. Among them, the kernel weight, mean matrix, and covariance matrix are all latent variables.
[0077] It can be understood that in one iteration process, after updating the latent variable, the value of the latent variable after this update can be used to update the responsiveness matrix in the next iteration process.
[0078] It can be seen that the calculation process of the responsiveness matrix corresponding to this iteration process is calculated based on the current value of the latent variable, and the sample weight is also considered. In this way, each responsiveness in the responsiveness matrix is actually a weighted responsiveness, which can not only reflect the current value of the latent variable but also reflect the sample weight of the corresponding training sample.
[0079] S4. Update the current value of the latent variable according to the responsiveness matrix corresponding to this iteration process;
[0080] In specific implementation, since the latent variables include kernel weights, mean matrices, and covariance matrices, the process of updating the current values of the latent variables in S4 may specifically include at least one of the following:
[0081] (1) For each Gaussian kernel, calculate the updated kernel weight of the Gaussian kernel according to the sample weights of each training sample and the responsiveness of each training sample to the Gaussian kernel in the responsiveness matrix corresponding to the current iteration process;
[0082] It can be understood that for each Gaussian kernel, the updated kernel weight of the Gaussian kernel can be calculated based on the sample weights of each training sample and the responsiveness of each training sample to the Gaussian kernel. In this way, the kernel weights of each Gaussian kernel can be obtained.
[0083] In some embodiments, the third calculation formula can be used to calculate the updated kernel weight of the k-th Gaussian kernel, and the third calculation formula includes:
[0084]
[0085] where, π k is the updated kernel weight of the k-th Gaussian kernel; r(z nk ) is the responsiveness of the n-th training sample to the k-th Gaussian kernel in the responsiveness matrix corresponding to the current iteration process; N is the number of training samples in the training sample set; W n is the sample weight of the n-th training sample.
[0086] It can be seen that for each Gaussian kernel, its updated kernel weight mainly depends on the sum of the responsiveness of each training sample to the Gaussian kernel in the responsiveness matrix corresponding to the current iteration process. In one iteration process, after updating the responsiveness matrix, the updated responsiveness matrix and the above third calculation formula can be used to update the kernel weights of each Gaussian kernel.
[0087] (2) For each Gaussian kernel, calculate the updated mean matrix of the Gaussian kernel according to the responsiveness of each training sample to the Gaussian kernel in the responsiveness matrix corresponding to the current iteration process and each training sample in the sample training set;
[0088] In specific implementation, the fourth calculation formula can be used to calculate the updated mean matrix of the k-th Gaussian kernel, and the fourth calculation formula includes:
[0089]
[0090] where, μ k is the updated mean matrix of the k-th Gaussian kernel, x n is the n-th training sample, r(znk ) is the responsiveness of the nth training sample in the responsiveness matrix corresponding to this iteration process to the kth Gaussian kernel, and N is the number of training samples in the training sample set.
[0091] That is to say, for each Gaussian kernel, the average value matrix can be calculated based on each training sample and the responsiveness of each training sample to this Gaussian kernel. In the fourth calculation formula, is the sum of the responsiveness of each training sample to the kth Gaussian kernel, and is the cumulative sum of the product of the responsiveness of each training sample to the kth Gaussian kernel and this training sample, which is equivalent to weighting each element in the training sample using the responsiveness. Finally, and The ratio of is used as the average value matrix, realizing the update of the average value matrix using the responsiveness matrix.
[0092] (3) For each Gaussian kernel, calculate the updated covariance matrix of this Gaussian kernel according to each training sample in the sample training set, the updated average value matrix of this Gaussian kernel, and the responsiveness of each training sample in the responsiveness matrix corresponding to this iteration process to this Gaussian kernel.
[0093] In specific implementation, the fifth calculation formula can be used in this step to calculate the updated covariance matrix of the kth Gaussian kernel, and the fifth calculation formula includes:
[0094]
[0095] In the formula, ∑ k is the updated covariance matrix of the kth Gaussian kernel, N is the number of training samples in the training sample set, x n is the nth training sample, r(z nk ) is the responsiveness of each training sample in the responsiveness matrix corresponding to this iteration process to this Gaussian kernel, μ k is the updated average value matrix of the kth Gaussian kernel. T is the inversion symbol of the vector.
[0096] That is to say, for each Gaussian kernel, the corresponding covariance matrix can be calculated using each training sample, the updated average value matrix of this Gaussian kernel, and the responsiveness of each training sample in the responsiveness matrix corresponding to this iteration process to this Gaussian kernel. It can be seen that the update of the covariance matrix is realized based on the previously updated average value matrix and responsiveness matrix.
[0097] In practice, regularization processing can be introduced when calculating the covariance matrix to prevent overfitting of the Gaussian mixture model in subsequent use.
[0098] The third, fourth and fifth calculation formulas are actually derived by maximizing the log-likelihood function. The updating process of the hidden variables using the three calculation formulas is actually the process of maximizing the calculation. Since S3 is essentially the process of expected value calculation and S4 is essentially the process of maximizing the calculation, it can be seen that the embodiment of the present invention actually solves the hidden variables through the expectation maximization (EM) algorithm. Compared with the traditional expectation maximization algorithm, the expectation maximization algorithm provided by the embodiment of the present invention takes into account the difference between the number of samples corresponding to different working conditions. The training samples under the working conditions with a small number of samples have a higher sample weight. In this way, after convergence, the kernel weights of each Gaussian kernel have a certain bias, and the standard solution results tend to give a local optimal result with a higher weight to the minority class samples. In this way, the working conditions with a small number of samples have a certain decision-making power in the subsequent fitting of the Gaussian mixture model, and will not be considered as abnormal classes due to too small a number, thereby improving the judgment of normal working conditions and potential risks in industrial scenarios.
[0099] S5, judging whether the current value of the updated latent variable meets the convergence condition;
[0100] If yes, then the method ends and the updated current value of the latent variable is taken as the optimal solution;
[0101] Otherwise, return to S3 to execute the next iteration process.
[0102] In specific implementation, the log-likelihood value can be used to determine whether the current value of the latent variable meets the convergence condition. For example, lnp(X|π,μ,∑) can be used to determine whether the current value of the latent variable meets the convergence condition, where π is the matrix formed by the kernel weights of each Gaussian kernel, μ is the high-dimensional matrix formed by the mean value matrix of each Gaussian kernel, and ∑ is the high-dimensional covariance matrix formed by the covariance matrix of each Gaussian kernel. If the change in the log-likelihood value between two adjacent iterations is less than a preset value, for example, 1e-4, it is considered to meet the convergence condition, otherwise it does not meet the convergence condition.
[0103] The latent variable determination method of the Gaussian mixture model provided in an embodiment of the present invention first initializes the latent variables, then determines the sample weight matrix according to the number of samples of each working condition, and then determines the responsiveness matrix corresponding to this iterative process according to the sample weight matrix and the latent variables, and then uses the responsiveness matrix of this iterative process to update the latent variables. If the updated latent variables meet the convergence conditions, the method is terminated and the current value of the updated latent variables is used as the optimal solution, that is, the current value of the latent variables at this time is the optimal result; otherwise, the next iterative process is entered until the optimal solution is obtained. In this scheme, the sample weights of each training sample are considered when calculating the responsiveness matrix, and the sample weights of each training sample under the condition with fewer samples are higher, that is, the difference between the number of samples corresponding to different working conditions is taken into account, and the training samples under the working condition with fewer samples have higher sample weights. In this way, after convergence, the kernel weights of each Gaussian kernel have a certain bias, and the standard solution results tend to give the minority class samples a higher weight of the local optimal result, so that the working condition with a small number of samples has a certain decision-making power in the subsequent fitting of the Gaussian mixture model, and will not be considered as an abnormal class because of the small number, thereby improving the judgment of normal working conditions and potential risks in industrial scenarios, that is, the latent variables calculated by this scheme can improve the early warning effect of the Gaussian mixture model. Moreover, this scheme does not need to increase the number of samples corresponding to the working condition by inserting artificial data into each sample under the working condition with a small number of samples, and will not introduce noise when the working conditions are relatively close. It can be seen that this scheme can better solve the problem of sample imbalance in industrial scenarios.
[0104] Moreover, compared with the traditional upsampling method, this scheme can avoid the uncertainty and potential risks introduced by artificial data. This scheme solves the problem of sample imbalance between different normal working conditions in industrial scenarios by increasing the weight of working conditions with a small number of samples through a latent variable calculation process with weight bias, so there is no need to upsample through simulated data. Compared with the traditional upsampling method, this scheme has lower requirements for the number of high-dimensional training samples, because the traditional upsampling method has a minimum data volume requirement for working conditions with a small number of samples, while this scheme emphasizes the representativeness of existing data. That is, when fitting the health condition characteristics of the equipment through the Gaussian mixture model, the quantity requirement becomes relatively low, which is especially suitable for the situation where the sample size is insufficient and representative for working conditions with a small number of samples. Therefore, this scheme is easier to promote.
[0105] In addition, for other solutions, for example, downsampling the training samples under the working conditions with a large number of samples, this downsampling method is often less robust and has a high application risk in the production working condition data with a more complex process. However, this solution will not reduce the robustness of the Gaussian mixture model and ensure the reliability of the Gaussian mixture model.
[0106] In a second aspect, an embodiment of the present invention provides an apparatus for determining latent variables of a Gaussian mixture model. Referring to Figure 2 , the apparatus 100 includes the following modules:
[0107] An initialization module 110, configured to set an initial value of the latent variable based on a training sample set of the Gaussian mixture model;
[0108] A weight determination module 120, configured to determine a sample weight matrix corresponding to the training sample set according to the number of samples corresponding to each working condition in the training sample set; wherein, the sample weight matrix includes the sample weights of each training sample in the training sample set; among the various working conditions corresponding to the training sample set, the fewer the number of samples corresponding to a working condition, the greater the sample weights of the respective training samples corresponding to this working condition;
[0109] A responsiveness determination module 130, configured to determine a responsiveness matrix corresponding to the current iteration process according to the sample weight matrix and the current value of the latent variable; wherein, the responsiveness matrix includes the responsiveness of each training sample under each working condition to each Gaussian kernel of the Gaussian mixture model, and the responsiveness of a training sample to a Gaussian kernel is the expected probability that the training sample belongs to the Gaussian kernel;
[0110] An update module 140, configured to update the current value of the latent variable according to the responsiveness matrix corresponding to the current iteration process;
[0111] A convergence judgment module 150, configured to judge whether the current value of the updated latent variable meets the convergence condition; if so, end the method provided by this apparatus, and use the current value of the updated latent variable as the optimal solution; otherwise, return to the responsiveness determination module to perform the next iteration process.
[0112] In one embodiment, the weight determination module 120 is specifically configured to: calculate the sample weight of the nth training sample in the training sample set by using a first calculation formula, and the first calculation formula includes:
[0113]
[0114] In the formula, W n is the sample weight of the nth training sample, s in is the number of samples corresponding to the ith working condition to which the nth training sample belongs, s j is the number of samples corresponding to the jth working condition, and I is the number of various working conditions corresponding to the training sample set.
[0115] In one embodiment, the latent variables include the kernel weights of each Gaussian kernel and the kernel parameters of each Gaussian kernel, and the kernel parameters include a mean matrix and a covariance matrix.
[0116] Further, the responsivity determination module 130 is specifically configured to: calculate the responsivity of the nth training sample in the responsivity matrix corresponding to the current iteration process with respect to the kth Gaussian kernel by using a second calculation formula, and the second calculation formula includes:
[0117]
[0118] wherein, r(z nk ) is the responsivity of the nth training sample in the responsivity matrix corresponding to the current iteration process with respect to the kth Gaussian kernel, W n is the sample weight of the nth training sample, π k is the current value of the kernel weight of the kth Gaussian kernel, x n is the nth training sample, μ k is the current mean matrix of the kth Gaussian kernel, ∑ k is the current covariance matrix of the kth Gaussian kernel, K is the number of Gaussian kernels, and N() is a Gaussian function.
[0119] In some embodiments, the update module 140 specifically includes:
[0120] A first calculation unit, configured to calculate the updated kernel weight of each Gaussian kernel according to the sample weights of each training sample and the responsivity of each training sample in the responsivity matrix corresponding to the current iteration process with respect to this Gaussian kernel;
[0121] A second calculation unit, configured to calculate the updated mean matrix of each Gaussian kernel according to the responsivity of each training sample in the responsivity matrix corresponding to the current iteration process with respect to this Gaussian kernel and each training sample in the sample training set;
[0122] A third calculation unit, configured to calculate the updated covariance matrix of each Gaussian kernel according to each training sample in the sample training set, the updated mean matrix of this Gaussian kernel, and the responsivity of each training sample in the responsivity matrix corresponding to the current iteration process with respect to this Gaussian kernel.
[0123] Further, the first calculation unit is specifically configured to: calculate the updated kernel weight of the kth Gaussian kernel by using a third calculation formula, and the third calculation formula includes:
[0124]
[0125] wherein, πk is the updated kernel weight of the k-th Gaussian kernel; r(z nk ) is the responsiveness of the n-th training sample in the responsiveness matrix corresponding to this iteration process to the k-th Gaussian kernel; N is the number of training samples in the training sample set; W n is the sample weight of the n-th training sample.
[0126] Further, the second calculation unit is specifically configured to: calculate the updated mean matrix of the k-th Gaussian kernel by using a fourth calculation formula, and the fourth calculation formula includes:
[0127]
[0128] In the formula, μ k is the updated mean matrix of the k-th Gaussian kernel, x n is the n-th training sample, r(z nk ) is the responsiveness of the n-th training sample in the responsiveness matrix corresponding to this iteration process to the k-th Gaussian kernel, and N is the number of training samples in the training sample set.
[0129] Further, the third calculation unit is specifically configured to: calculate the updated covariance matrix of the k-th Gaussian kernel by using a fifth calculation formula, and the fifth calculation formula includes:
[0130]
[0131] In the formula, ∑ k is the updated covariance matrix of the k-th Gaussian kernel, N is the number of training samples in the training sample set, x n is the n-th training sample, r(z nk ) is the responsiveness of each training sample in the responsiveness matrix corresponding to this iteration process to this Gaussian kernel, and μ k is the updated mean matrix of the k-th Gaussian kernel.
[0132] It can be understood that the explanations, specific implementation manners, beneficial effects, examples, etc. of the relevant content in the device provided in the embodiments of the present invention can refer to the corresponding parts in the method provided in the first aspect, and will not be elaborated here.
[0133] In a third aspect, an embodiment of the present invention provides a computing device, and the device includes: at least one memory and at least one processor;
[0134] The at least one memory is used to store a machine-readable program;
[0135] The at least one processor is used to call the machine-readable program and execute the method provided in the first aspect.
[0136] It is understandable that for the explanations, specific implementation manners, beneficial effects, examples, etc. of the relevant content in the device provided in the embodiments of the present invention, reference may be made to the corresponding parts in the method provided in the first aspect, and details are not described herein again.
[0137] Fourthly, an embodiment of the present invention provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor is caused to execute the method provided in the first aspect.
[0138] Specifically, a system or device equipped with a storage medium may be provided, on which software program codes for implementing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program codes stored in the storage medium.
[0139] In this case, the program codes read from the storage medium itself can implement the functions of any one of the above embodiments, so the program codes and the storage medium storing the program codes constitute a part of the present invention.
[0140] Embodiments of the storage medium for providing program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program codes may be downloaded from a server computer via a communication network.
[0141] In addition, it should be clear that not only can the actual operations be completed in part or in whole by executing the program codes read by the computer, but also by means of instructions based on the program codes, the operating system operating on the computer, etc., so as to implement the functions of any one of the above embodiments.
[0142] In addition, it can be understood that the program codes read from the storage medium are written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion module connected to the computer, and then based on the instructions of the program codes, the CPU, etc. installed on the expansion board or the expansion module are caused to execute part and all of the actual operations, so as to implement the functions of any one of the above embodiments.
[0143] It is understandable that for the explanations, specific implementation manners, beneficial effects, examples, etc. of the relevant content in the computer-readable medium provided in the embodiments of the present invention, reference may be made to the corresponding parts in the method provided in the first aspect, and details are not described herein again.
[0144] Each embodiment in this specification is described in a progressive manner. For the identical or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the description of the method embodiments.
[0145] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, add-ons, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0146] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solution of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for determining latent variables of a Gaussian mixture model, characterized in that, Including: S1. Set the initial value of the latent variable based on the training sample set of the Gaussian mixture model; S2. Determine the sample weight matrix corresponding to the training sample set according to the number of samples corresponding to each working condition in the training sample set; wherein, the sample weight matrix includes the sample weights of each training sample in the training sample set; among the various working conditions corresponding to the training sample set, the fewer the number of samples corresponding to a working condition, the greater the sample weights of each training sample corresponding to this working condition; S3. Determine the responsiveness matrix corresponding to the current iteration process according to the sample weight matrix and the current value of the latent variable; wherein, the responsiveness matrix includes the responsiveness of each training sample under each working condition to each Gaussian kernel of the Gaussian mixture model, and the responsiveness of a training sample to a Gaussian kernel is the expected probability that this training sample belongs to this Gaussian kernel; S4. Update the current value of the latent variable according to the responsiveness matrix corresponding to the current iteration process; S5. Determine whether the current value of the updated latent variable meets the convergence condition; If so, end this method and use the current value of the updated latent variable as the optimal solution; Otherwise, return to S3 to execute the next iteration process, wherein, the latent variable includes the kernel weight of each Gaussian kernel and the kernel parameters of each Gaussian kernel, and the kernel parameters include the mean matrix and the covariance matrix, wherein, determining the responsiveness matrix corresponding to the current iteration process according to the sample weight matrix and the current value of the latent variable includes: Calculating the responsiveness of the nth training sample in the current iteration process to the kth Gaussian kernel in the responsiveness matrix using the second calculation formula, and the second calculation formula includes: where r(z nk ) is the responsiveness of the n-th training sample in the responsiveness matrix corresponding to this iteration process to the k-th Gaussian kernel, W n is the sample weight of the n-th training sample, π k is the current value of the kernel weight of the k-th Gaussian kernel, x n is the n-th training sample, μ k is the current mean matrix of the k-th Gaussian kernel, ∑ k is the current covariance matrix of the k-th Gaussian kernel, K is the number of Gaussian kernels, N() is the Gaussian function, wherein, updating the current value of the latent variable according to the responsiveness matrix corresponding to the current iteration process includes at least one of the following: For each Gaussian kernel, calculate the updated kernel weight of this Gaussian kernel according to the sample weights of each training sample and the responsiveness of each training sample in the responsiveness matrix corresponding to the current iteration process to this Gaussian kernel; For each Gaussian kernel, calculate the updated mean matrix of this Gaussian kernel according to the responsiveness of each training sample in the responsiveness matrix corresponding to the current iteration process to this Gaussian kernel and each training sample in the sample training set; For each Gaussian kernel, calculate the updated covariance matrix of this Gaussian kernel according to each training sample in the sample training set, the updated mean matrix of this Gaussian kernel, and the responsiveness of each training sample in the responsiveness matrix corresponding to the current iteration process to this Gaussian kernel.
2. The method according to claim 1, characterized in that, Determining the sample weight matrix corresponding to the training sample set according to the number of samples corresponding to each working condition in the training sample set includes: Calculating the sample weight of the nth training sample in the training sample set using the first calculation formula, and the first calculation formula includes: where W n is the sample weight of the nth training sample, s in is the number of samples corresponding to the ith working condition to which the nth training sample belongs, s j is the number of samples corresponding to the jth working condition, and I is the number of each working condition corresponding to the training sample set.
3. The method according to claim 1, wherein For each Gaussian kernel, calculating the updated kernel weight of the Gaussian kernel according to the sample weights of each training sample and the responsiveness of each training sample in the responsiveness matrix corresponding to the current iteration process to the Gaussian kernel, including: Calculating the updated kernel weight of the k-th Gaussian kernel by using a third calculation formula, where the third calculation formula includes: where, π k is the updated kernel weight of the k-th Gaussian kernel; r(z nk ) is the responsiveness of the n-th training sample in the responsiveness matrix corresponding to the current iteration process to the k-th Gaussian kernel; N is the number of training samples in the training sample set; W n is the sample weight of the n-th training sample.
4. The method according to claim 1, characterized in that For each Gaussian kernel, calculating the updated mean matrix of the Gaussian kernel according to the responsiveness of each training sample in the responsiveness matrix corresponding to the current iteration process to the Gaussian kernel and each training sample in the sample training set, including: Calculating the updated mean matrix of the k-th Gaussian kernel by using a fourth calculation formula, where the fourth calculation formula includes: where μ k is the updated mean matrix of the k-th Gaussian kernel, x n is the n-th training sample, r(z nk ) is the responsiveness of the n-th training sample in the responsiveness matrix corresponding to the current iteration process to the k-th Gaussian kernel, and N is the number of training samples in the training sample set.
5. The method according to claim 1, wherein For each Gaussian kernel, calculating the updated covariance matrix of the Gaussian kernel according to each training sample in the sample training set, the updated mean matrix of the Gaussian kernel, and the responsiveness of each training sample in the responsiveness matrix corresponding to the current iteration process to the Gaussian kernel, including: Calculating the updated covariance matrix of the k-th Gaussian kernel by using a fifth calculation formula, where the fifth calculation formula includes: Where, ∑k is the covariance matrix of the k-th Gaussian kernel after update, N is the number of training samples in the training sample set, x n is the n-th training sample, r(z nk ) is the responsiveness of each training sample in the responsiveness matrix corresponding to this iteration process to this Gaussian kernel, μ k is the mean matrix of the k-th Gaussian kernel after update.
6. An apparatus for determining hidden variables of a Gaussian mixture model, characterized in that, The device includes: An initialization module, configured to set an initial value of the latent variable based on a training sample set of the Gaussian mixture model; A weight determination module, configured to determine a sample weight matrix corresponding to the training sample set according to the number of samples corresponding to each working condition in the training sample set; where the sample weight matrix includes the sample weights of each training sample in the training sample set; among the working conditions corresponding to the training sample set, the fewer the number of samples corresponding to a working condition, the greater the sample weights of each training sample corresponding to the working condition; A responsiveness determination module, configured to determine a responsiveness matrix corresponding to the current iteration process according to the sample weight matrix and the current value of the latent variable; where the responsiveness matrix includes the responsiveness of each training sample under each working condition to each Gaussian kernel of the Gaussian mixture model, and the responsiveness of a training sample to a Gaussian kernel is the expected probability that the training sample belongs to the Gaussian kernel; An update module, configured to update the current value of the latent variable according to the responsiveness matrix corresponding to the current iteration process; A convergence determination module, configured to determine whether the current value of the updated latent variable satisfies a convergence condition; if so, ending the method provided by the device and using the current value of the updated latent variable as an optimal solution; otherwise, returning to the responsiveness determination module to perform the next iteration process, where the latent variable includes the kernel weight of each Gaussian kernel and the kernel parameters of each Gaussian kernel, and the kernel parameters include a mean matrix and a covariance matrix, where determining the responsiveness matrix corresponding to the current iteration process according to the sample weight matrix and the current value of the latent variable includes: Calculating the responsiveness of the n-th training sample in the responsiveness matrix corresponding to the current iteration process to the k-th Gaussian kernel by using a second calculation formula, where the second calculation formula includes: where r(z nk ) is the responsiveness of the nth training sample in the responsiveness matrix corresponding to this iteration process for the kth Gaussian kernel, W n is the sample weight of the nth training sample, π k is the current value of the kernel weight of the kth Gaussian kernel, x n is the nth training sample, μ k is the current mean matrix of the kth Gaussian kernel, ∑ k is the current covariance matrix of the kth Gaussian kernel, K is the number of Gaussian kernels, N() is the Gaussian function, Among them, updating the current value of the latent variable according to the responsiveness matrix corresponding to the current iteration process includes at least one of the following: For each Gaussian kernel, calculate the updated kernel weight of the Gaussian kernel according to the sample weights of each training sample and the responsiveness of each training sample to the Gaussian kernel in the responsiveness matrix corresponding to the current iteration process; For each Gaussian kernel, calculate the updated mean matrix of the Gaussian kernel according to the responsiveness of each training sample to the Gaussian kernel in the responsiveness matrix corresponding to the current iteration process and each training sample in the sample training set; For each Gaussian kernel, calculate the updated covariance matrix of the Gaussian kernel according to each training sample in the sample training set, the updated mean matrix of the Gaussian kernel, and the responsiveness of each training sample to the Gaussian kernel in the responsiveness matrix corresponding to the current iteration process.
7. A computing device, characterized in that, The device includes: at least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program and execute the method according to any one of claims 1 to 5.
8. A computer-readable medium, characterized in that, Computer instructions are stored on the computer-readable medium, and when the computer instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for training hybrid model
CN107292323A
Self-adaptive soft measurement method based on semi-supervised incremental Gaussian mixture regression
CN112650063A