Method and system for guaranteeing local dx-private by combining bounded disturbance generation mechanism with TKmeans algorithm
Through the combination of bounded perturbation generation mechanism and TKmeans algorithm, the balance problem of LDP method between privacy protection and data clustering quality is solved, and efficient clustering is achieved under high-dimensional data and uncertain distribution is improved, thus improving clustering quality and stability.
Patent Information
- Application Number
- CN202510391635.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The existing local differential privacy (LDP) method is difficult to achieve in the balance between privacy protection and data clustering quality, especially when dealing with high-dimensional data and uncertain data distribution, it is often difficult to obtain better clustering quality.
The bounded perturbation generation mechanism (BPGM) is used to combine the TKmeans algorithm to perturb the data as a whole on the user side, and the K-means algorithm is improved on the server side to ensure dx-privacy and improve the clustering quality.
On the premise of ensuring user privacy, the quality and stability of data clustering are improved, different data distributions are adapted to, and privacy protection and data utility are balanced.
Smart Images

Figure SMS_18 
Figure SMS_20 
Figure SMS_21
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data mining, and relates to a synthetic data generation mechanism and a clustering algorithm, in particular to a bounded perturbation generation mechanism and a TKmeans clustering method. Background Art
[0002] With the development of the big data era, data analysis technology has become a key technology for various service providers to analyze user behavior and improve service levels. As one of the most popular unsupervised learning algorithms, K-means clustering has been widely applied in many data-driven applications, such as market analysis, anomaly detection, recommendation systems, and key industrial sectors such as manufacturing, transportation and logistics, energy, and healthcare. Due to the enactment of data privacy regulations in many jurisdictions, the task of performing effective data analysis has become increasingly challenging. Therefore, how to effectively perform data analysis while protecting privacy has become an important issue.
[0003] In this context, local differential privacy (LDP) has been widely used by many large technology companies due to its strict privacy guarantee. Google has deployed LDP to the Chrome browser to collect user data for analyzing their preferences while protecting user privacy; Microsoft and Apple use randomized response to collect and analyze user information to meet differential privacy requirements. In addition, Google's TensorFlow library meets the requirements of differential privacy when building deep learning models; IBM has developed the diffprivlib library to provide a reference for differential privacy research. The advantage of LDP is that data containing user privacy is perturbed locally by the user and uploaded to the server, and the server cannot restore the original data, thus achieving strong privacy protection.
[0004] Benefiting from the advantages of local differential privacy, researchers have begun to explore how to perform Kmeans clustering analysis while protecting user privacy. The EUGkM model proposed by Su et al. combines interactive and non-interactive strategies, improving the clustering quality. However, the clustering quality of this method decreases significantly and the computational cost is high when the data dimension is high or the distribution is uneven. Xia et al. transformed the user feature vectors into binary strings for random perturbation to improve the clustering accuracy, but it increases the computational and communication costs. In LDP, the distinguishability between any pair of data records is the same. Intuitively, this provides the same level of privacy guarantee for users' sensitive data as any other data. A characteristic of LDP is that the cumulative difference between different reports is relatively large. Therefore, recent research has not been limited to data value perturbation, but focuses on the overall properties of data. For example, the SamPrivSyn method proposed by Chen et al. guarantees local differential privacy by constructing a synthetic dataset based on two-dimensional marginal distributions, but it has a high computational complexity and the generated synthetic data is difficult to interpret. The DPM clustering algorithm proposed by Schütt et al. projects data points into one-dimensional space and uses the exponential mechanism to select the optimal window, but it encounters bottlenecks in high-dimensional datasets. Especially when the data contains noise or outliers, it is difficult to find a suitable separator, and the obvious separation boundaries assumed often do not hold.
[0005] Generally speaking, it is still difficult to balance privacy protection and data utility in current local differential privacy (LDP) methods. Especially when dealing with high-dimensional data and uncertain data distributions, it is often difficult to obtain good clustering quality. To effectively solve this problem and find a suitable balance between privacy protection and data clustering quality, we introduce the definition of generalized differential privacy, namely d x -privacy.
[0006] d x The definition of d x -privacy provides privacy guarantees for data with a distance function d(x) in a metric space. This distance can be the Manhattan distance, Euclidean distance, and Chebyshev distance. d -privacy requires that the distinguishability level of two data is based on the distance between two data records. In the clustering algorithm, the data space is defined as a d-dimensional real space x which contains the Euclidean distance (i.e., L2 distance). Therefore, it is very desirable to perturb the user data as a whole rather than perturbing each dimension independently to ensure d Summary of the Invention
[0007] In view of the deficiencies of the prior art, the present invention provides a method and system for ensuring local differential privacy by combining a Bounded Perturbation Generation Mechanism (BPGM) with the TKmeans algorithm, which can guarantee the differential privacy of users. At the user side, BPGM perturbs the data as a whole instead of processing each dimension independently. To ensure the interpretability of the generated data, we set the noise distance of the output of BPGM within a bounded domain and set an upper limit l, which can be selected according to the user's privacy protection level. We also hope that the distribution of the output value is positively correlated with its distance from the actual value. Specifically, we design a probability density function of a bounded exponential distribution, which decreases when the distance from the actual value increases from 0 to a predetermined public constant l. Users can use the gradient descent method to generate synthetic data for the original data based on the sampled noise distance. To eliminate the impact of outliers and heavy-tailed data that may be brought by the synthetic data set on the clustering quality and to adapt to data with different data distributions, we also improve the K-means algorithm using the T mixture distribution combined with the Expectation-Maximization algorithm (EM) on the server side, effectively solving the balance problem between data utility and clustering quality. x -privacy. At the user side, BPGM perturbs the data as a whole instead of processing each dimension independently. x -privacy. Specifically, at the user side, BPGM perturbs the data as a whole instead of processing each dimension independently. To ensure the interpretability of the generated data, we set the noise distance of the output of BPGM within a bounded domain and set an upper limit l, which can be selected according to the user's privacy protection level. We also hope that the distribution of the output value is positively correlated with its distance from the actual value. Specifically, we design a probability density function of a bounded exponential distribution, which decreases when the distance from the actual value increases from 0 to a predetermined public constant l. Users can use the gradient descent method to generate synthetic data for the original data based on the sampled noise distance. To eliminate the impact of outliers and heavy-tailed data that may be brought by the synthetic data set on the clustering quality and to adapt to data with different data distributions, we also improve the K-means algorithm using the T mixture distribution combined with the Expectation-Maximization algorithm (EM) on the server side, effectively solving the balance problem between data utility and clustering quality.
[0008] Therefore, the specific technical solution of the present invention at the user side is as follows:
[0009] Step 1: Preprocess the sample data and normalize the data;
[0010] Step 2: Perform inverse transform sampling according to the bounded exponential function to obtain the noise distance
[0011] Step 3: Randomly initialize the synthetic data Then combine the distance between the real data r and the synthetic data with the noise distance sampled in Step 2 to define the loss function, and finally use the gradient descent method to continuously update the synthetic data record;
[0012] Furthermore, in Step 1, the data preprocessing method is MaxMin normalization, which is a linear normalization that linearly maps the original data so that the data is mapped between [0,1]. Through the normalization formula the data is normalized to the space [0,1].
[0013] Furthermore, in Step 2, perform inverse transform sampling according to the bounded exponential function to obtain the noise distance so that any possible data record and the distance between the user's real data record Normalization factor Ensures that the sampled noise distance is within [0, l], where l is a constant specified by the user, representing the upper bound of the sampled noise distance. Different values of l will sample different noise distances under the same ∈.
[0014] Furthermore, in step 3, a batch of synthetic data is randomly initialized. The distance between the real data r and the synthetic data and the sampled noise distance is defined as the loss function and the process of obtaining the final synthetic data is modeled as an optimization process, and the synthetic data record is continuously updated using the gradient descent method until it meets a certain stopping condition.
[0015] The specific technical solution of the present invention on the server side is as follows:
[0016] Step 1: Construct a complete data set and a log-likelihood function;
[0017] Step 2: Calculate the expected value based on the E-step of the EM algorithm to estimate the hidden variables;
[0018] Step 3: Maximize the log-likelihood function based on the M-step of the EM algorithm to update each parameter;
[0019] Step 4: Repeat steps 2 to 3 until the Q function converges or reaches the maximum number of iterations.
[0020] Furthermore, in step 1, we assume that the data is sampled from a T-mixture distribution, and the probability density function is given by the following formula:
[0021]
[0022] The data preconditioning of the TKmeans algorithm is to take into account the latent variable z in the mixture model (TMM) nk (indicating whether the sample x n belongs to the k-th cluster) and the missing data u n (representing the scale parameter of the sample x n ). Therefore, the complete data is:
[0023]
[0024] Next, the distribution satisfied among the data is:
[0025]
[0026] Therefore, the log-likelihood function of the complete data can be expressed as:
[0027] ln Lc (Ψ|x,u,z) = ln L G (v|u,z) + ln L N (μ,α|x,u,z),
[0028] In the EM algorithm, the objective function of the new iteration is the current conditional expectation of the complete-data log-likelihood, where the parameters marked with * need to be updated in each iteration. The objective function is:
[0029] Q(Ψ * |Ψ) = E q (ln L c (Ψ|x,u,z)) = Q1(v * |Ψ) + Q2(μ * ,α * |Ψ), Ψ = {π,ν,μ,∑},
[0030]
[0031] Furthermore, in step 2, the expected value is calculated according to the E-step in the EM algorithm to estimate the hidden variables.
[0032] First, calculate the posterior probability E(z nk |x n ), that is, the posterior probability of each sample belonging to each cluster, which is:
[0033]
[0034] Then calculate the expectation of the missing data E(u n |x n ,z n ). Since and u n (x n -μ k ) T (x n -μ k ) follows distribution, according to the properties of the gamma distribution, the possibility of u n is Then, given x n ,z nk = 1, the posterior distribution of u n is The expectation of the missing value u n is:
[0035]
[0036] Then calculate E(lnu n |xn , z n ), according to the lemma: If the random variable R ~ gamma(a, b), then E(lnR) = φ(a) - lnb, It is obtained that:
[0037]
[0038] Furthermore, in step 3, according to the M-step of the EM algorithm to maximize the log-likelihood function, each parameter is updated.
[0039] First, according to Q2(μ * , α * |Ψ) and the posterior probabilities u nk and τ nk obtained in the E-step, update the parameters μ * , α * , that is:
[0040]
[0041] Then, from Q1(v * |Ψ) and the posterior probabilities obtained in E, it can be known that updating v * is to obtain the approximate solution of the following equation, that is:
[0042]
[0043] where η is a constant, which can be calculated from τ nk , u nk obtained in the E-step, that is:
[0044]
[0045] To find the solution of the equation, we note that the theorem B i is the second Bernoulli number and Applying this theorem to the original equation, the original equation can be transformed into: Finally, the solution is obtained as:
[0046]
[0047] Compared with the prior art, the advantages of the present invention are:
[0048] (1) Using a bounded perturbation generation mechanism to ensure local d x -privacy. In the clustering algorithm, the data space is defined as a d-dimensional real space which includes the Euclidean distance (i.e., the L2 distance). Therefore, using d x-privacy has an advantage over traditional LDP in balancing the degree of privacy protection and clustering quality.
[0049] (2) TKmeans is adopted to eliminate the problems of outliers and heavy-tailed distributions brought by synthetic data generated by different client sides, better adapt to various different data sets, and further improve the clustering quality and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the specific technical solutions implemented by the present invention, the following will briefly introduce the drawings required in the specific implementation manners or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0051] Figure 1 It is a schematic diagram of the system model described in the present invention;
[0052] Figure 2 It is a schematic diagram of noise distance sampling and generating synthetic data;
[0053] Figure 3 It is the noise distance probability distribution of the bounded perturbation generation mechanism (BPGM) with different ∈ at l = 1;
[0054] Figure 4 It is the noise distance probability distribution of the bounded perturbation generation mechanism (BPGM) with different ∈ at l = 10.
[0055] Figure 5 It is the performance of BPGMTKl = * and the traditional LDP method on different data sets, with ARI as the y-axis and ∈ as the x-axis. (a) Iris, (b) Seeds, (c) Wine, (d) BuddyMove.
[0056] Figure 6 It is the performance of BPGMTKl = * and the traditional LDP method on different data sets, with RE as the y-axis and ∈ as the x-axis. (a) Iris, (b) Seeds, (c) Wine, (d) BuddyMove. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0058] A bounded perturbation mechanism combined with the TKmeans clustering algorithm to protect local d x -privacy method and the model diagram of the system are as follows Figure 1 As shown. The model contains two entities: (1) User side: Provides the user's original data, and a perturbator is deployed on the user side. The perturbator executes the bounded perturbation generation mechanism (BPGM) method to perturb the user data and finally obtains the perturbed data and sends the perturbed data to the server side; (2) Server side: Responsible for receiving the perturbed data from each user side Collects the data of all users and directly uses the TKmeans algorithm to perform tasks such as clustering analysis on it.
[0059] First, the following steps are included on the user side in sequence:
[0060] Step 1: Preprocess the sample data and normalize the data;
[0061] (1) For each user, normalize each collected user data record through the normalization formula Normalize the data to the space [0, 1].
[0062] Step 2: Obtain the noise distance through inverse transform sampling according to the bounded exponential function
[0063] (1) To obtain the noise distance First, simplify the distance function to a scalar and let
[0064] (2) Then integrate F BPGM to obtain the cumulative distribution function
[0065] (3) Input the random variable u ∼ U(0, 1) into the inverse cumulative distribution function CDF -1 (u), and solve the equation through the inverse function to generate the sampling value That is
[0066] Step 3: Randomly initialize the synthetic data Then combine the distance between the real data r and the synthetic data with the noise distance sampled in step two Define it as the loss function, and finally use the gradient descent method to continuously update the synthetic data record.
[0067] (1) Randomly initialize a batch of synthetic data Define the distance between the real data \(r\) and the synthetic data and the noise distance obtained by sampling as the loss function
[0068] (2) Expand the square of the loss function For any calculate the gradient Therefore, the gradient of the squared loss is
[0069] (3) Then update the synthetic data using the following update formula
[0070] (4) Repeat steps (1), (2), and (3) until the loss function \(loss\) converges to a small value or reaches the maximum number of iterations, then terminate.
[0071] After the user's original data is processed by BPGM, it is sent to the server. The server is responsible for collecting the perturbation data of each user and directly clustering it using the TKmeans algorithm.
[0072] The steps on the server are as follows:
[0073] Step 1: Construct the complete dataset and the log-likelihood function;
[0074] (1) Construct the complete dataset. We assume that the data is sampled from a \(T\)-mixture distribution. Let the dataset where \(\pi=\{\pi k |k = 1,\ldots,K\}\) and \(\mu=\{\mu k |k = 1,\ldots,k\}\) is the mean vector and \(\sum k =\{\sum k |k = 1,\ldots,K\}\) is the covariance matrix and \(\sum k \in\prod(p)\). The probability density function is given by the following formula:[[]]
[0075]
[0076] where \(\Psi=\{\pi,v,\mu,\sum\}\), \(v = \{v k |k = 1,\ldots,K\}\).
[0077] The data preconditioning of the TKmeans algorithm is to make the latent variable \(z\) in the mixture model (TMM) nk (indicating whether the sample \(x n belongs to the \(k\)-th cluster) and the missing data \(u\)n (The scale parameter representing the sample x n ) is taken into account, so the complete data is:
[0078]
[0079] where z1,…,z n are the indicators related to clustering, u1,…,u N , n = {1...N} are the data records with missing Gaussian distributions.
[0080] (2) Construct the log-likelihood function. Let I denote the identity matrix with the dimension of the sample dimension and ∑ k = αI, k = (1,...K). In addition, we also assume that v k = v, k = (1,…K), which is an indicator related to clustering and can adjust the complexity of the model.
[0081] First, we give the distributions satisfied among the various data:
[0082]
[0083]
[0084] Therefore, the log-likelihood function of the complete data can be expressed as:
[0085] ln L c (Ψ|x,u,z) = ln L G (v|u,z) + ln L N (μ,α|x,u,z).
[0086] In the EM algorithm, the objective function of the new iteration is the current conditional expectation of the log-likelihood of the complete data, where the parameters marked with * are those to be updated in each iteration. The objective function is:
[0087] Q(Ψ * |Ψ) = E q (ln L c (Ψ|x,u,z)) = Q1(v * |Ψ) + Q2(μ * ,α * |ψ), ψ = {π,v,μ,Σ},
[0088]
[0089] Step 2: Calculate the expected value based on the E-step of the EM algorithm to estimate the latent variables;
[0090] (1) Calculate the posterior probability E(z nk |x n), that is, the posterior probability of each sample belonging to each cluster, namely:
[0091]
[0092] where t k (x n |v, μ k , αI) represents the probability density function of the t-distribution, v is the degrees of freedom parameter, μ k is the mean vector, and αI is the covariance matrix.
[0093] (2) Calculate the expectation E(u n |x n , z n ). Since and u n (x n - μ k ) T (x n - μ k ) follows distribution, the possibility of u n can be obtained according to the properties of the gamma distribution as Then, given x n , z nk = 1, the posterior distribution of u n is The expectation of the missing value u n is:
[0094]
[0095] (3) Calculate E(ln u n |x n , z n ). According to the lemma: If the random variable R ~ gamma(a, b), then we get:
[0096]
[0097] Step 3: Maximize the log-likelihood function based on the M-step of the EM algorithm and update each parameter;
[0098] (1) Update the parameters μ * , α * |Ψ) according to Q2(μ nk and τ nk obtained in the E-step and the posterior probabilities u * , α * , that is:
[0099]
[0100] (2) According to the posterior probability obtained from Q1(v * |Ψ) and E, it can be known that updating v * is to obtain the approximate solution of the following equation, that is:
[0101]
[0102] where η is a constant, which can be calculated from τ nk , u nk obtained in the E-step, that is:
[0103]
[0104] To find the solution of the equation, we note that the theorem B i is the second kind of Bernoulli number and Applying this theorem to the original equation, the original equation can be transformed into: Finally, the solution is obtained as:
[0105]
[0106] Step 4: Repeat Step 2 to Step 3 until the Q function converges or reaches the maximum number of iterations.
[0107] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Those of ordinary skill in the art of this technology, within the scope of the essence of the present invention, any changes, modifications, additions or substitutions should fall within the protection scope of the present invention.
Claims
1. Bounded perturbation generation mechanism, characterized in that: Distinguish the user's real data record from the perturbed data record through a one-dimensional distance attribute. If the attacker cannot know how far the user's original data record is from the reported data record, then it cannot succeed in the attack, and thus this method achieves d x -privacy. The method includes the following steps: Step 1: Preprocess the sample data and normalize the data; Step 2: According to the bounded exponential function perform inverse transform sampling to obtain the noise distance Step 3: Randomly initialize synthetic data Next, combine the distance between the real data r and the synthetic data with the noise distance sampled in Step 2 to define the loss function, and finally use the gradient descent method to continuously update the synthetic data record.
2. The bounded disturbance generation mechanism according to claim 1, wherein: In Step 1, the data preprocessing method is MaxMin normalization, which is a linear normalization that linearly maps the original data so that the data is mapped between [0, 1]. Through the normalization formula the data is normalized to the space [0, 1].
3. The bounded disturbance generation mechanism according to claim 1, wherein: In step 2, according to the bounded exponential function perform inverse transform sampling to obtain the noise distance so that the distance between any possible data record and the user's true data record normalization factor ensures that the sampled noise distance is in [0, l], where l is a constant specified by the user, representing the upper bound of the sampled noise distance. Different values of l will sample different noise distances under the same ∈; to obtain the noise distance first simplify the distance function to a scalar and let then integrate F BPGM to obtain the cumulative distribution function input the random variable u ∼ U(0, 1) into the inverse cumulative distribution function CDF -1 (u), and solve the equation through the inverse function to generate the sampled value that is 4. The bounded perturbation generation mechanism according to claim 1, characterized in that: In step 3, a batch of synthetic data is randomly initialized The distance between the real data r and the synthetic data is defined as the loss function with the noise distance obtained by sampling And the process of obtaining the final synthetic data is modeled as an optimal process, and the synthetic data record is continuously updated using the gradient descent method until it meets a certain stopping condition; specifically, first expand the square of the loss function For any Take the gradient Therefore, the gradient of the squared loss is The update formula is Then continuously iterate the above process until the loss function loss converges to a small value or reaches the maximum number of iterations, and then terminate. 5. The TKmeans clustering algorithm, characterized in that: Derive the TKmeans algorithm using the T mixture model (TMM) combined with the expectation maximization (EM) algorithm. The T mixture distribution is a distribution with heavy-tailed characteristics, and its tail is thicker, which can better handle outliers. Since the T mixture distribution is insensitive to outliers, our algorithm performs more robustly when dealing with datasets containing outliers. Moreover, since the algorithm utilizes the information of all samples, its stability is also improved. The method includes the following steps: Step 1: Construct the complete dataset and the log-likelihood function; Step 2: Calculate the expected value based on the E-step of the EM algorithm and estimate the hidden variables; Step 3: Maximize the log-likelihood function based on the M-step of the EM algorithm and update each parameter; Step 4: Repeat Step 2 to Step 3 until the Q function converges or reaches the maximum number of iterations.
6. The TKmeans clustering algorithm described in claim 5, wherein: In step 1, we assume that the data is sampled from a T mixture distribution, and let the data set where π = {π k | k = 1, …, K} and μ = {μ k | k = 1, …, K} is the mean vector and ∑ k = {∑ k | k = 1, …, K} is the covariance matrix and Σ k ∈ ∏(p), and the probability density function is given by the following formula: where ψ = {π, v, μ, ∑}, v = {v k | k = 1, …, K}; The data pre - processing of the TKmeans algorithm is to include the latent variable z in the T - mixture model (TMM) nk (indicating whether the sample x n belongs to the k - th cluster) and the missing data u n (representing the scale parameter of the sample x n ). Therefore, the complete data is as follows: where z1, …, z n are clustering-related metrics, u1, …, u N , and n = {1...N} are the data records with missing Gaussian distributions; To construct a log-likelihood function that includes complete data, we let I denote the identity matrix of dimension equal to the sample dimension and ∑ k = αI, k = (1,...K). In addition, we also assume that v k = v, k = (1,…K), which is an index related to clustering and can adjust the complexity of the model; First, we give the distributions satisfied among the various data: Therefore, the log-likelihood function of the complete data can be expressed as: ln L c (Ψ|x,u,z) = ln L G (v|u,z) + ln L N (μ,α|x,u,z); In the EM algorithm, the objective function of the new iteration is the current conditional expectation of the complete data log-likelihood, where the parameters with * are the ones to be updated in each iteration. The objective function is: Q(Ψ * |Ψ) = E q (ln L c (Ψ|x,u,z)) = Q1(v * |Ψ)+Q2(μ * ,α * |Ψ), Ψ = {π,v,μ,∑}; 7. The TKmeans clustering algorithm according to claim 5, wherein In Step 2: Calculate the expected value and estimate the hidden variables according to the E-step in the EM algorithm; First, calculate the posterior probability E(z nk |x n ), that is, the posterior probability that each sample belongs to each cluster, namely: where t k (xx|ν,μ k ,αI) represents the probability density function of the t-distribution, ν is the degrees of freedom parameter, μ k is the mean vector, and αI is the covariance matrix; Next, calculate the expectation E(u n |x n ,z n ). Since and u n (x n -μ k ) T (x n -μ k ) follows a distribution, according to the properties of the gamma distribution, the probability of u n can be obtained as Then, given x n ,z nk = 1, the posterior distribution of u n is The missing value u n has an expectation of: Next, calculate E(lnu n |x n ,z n ). According to the lemma: if the random variable R ~ gamma(a, b), then E(lnR) = φ(a) - lnb, We get:
8. The TKmeans clustering algorithm according to claim 5, characterized in that, In Step 3: Maximize the log-likelihood function according to the M-step in the EM algorithm; First, according to Q2(μ ★ ,α ★ |Ψ) and the posterior probabilities u nk and τ nk obtained in the E-step, update the parameters μ ★ ,α ★ , that is: Then, from the posterior probability obtained from Q1(v ★ |ψ) and E, it can be known that updating v ★ is to obtain an approximate solution to the following equation, that is: where η is a constant that can be calculated from τ obtained in the E-step nk , u nk Calculated as follows: To find the solution of the equation, we note that the theorem B i is the second kind of Bernoulli number and Applying this theorem to the original equation, the original equation can be transformed into: Finally, the solution is obtained as: