Training method and device for bag-of-words processing model, bag-of-words processing method and device

By using EM algorithm and variational inference to optimize the posterior distribution model in the training of bag-of-word processing model, the problem of slow convergence speed and high noise influence is solved, and faster convergence and higher robustness are achieved.

CN114969327BActive Publication Date: 2025-05-06ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210447095.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-05-06
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

During the training process of bag-of-word processing model, the model convergence speed is slow, and the existence of noise variables and noise labels seriously affect the training process and generalization performance of the model.

Method used

The EM algorithm-based training method is used to train a bag-of-word processing model containing multiple hidden variables. In step E, the posterior distribution model is iteratively optimized using the update rules of variational inference and the fixed point, and the expectations of the hidden variables are calculated. In step M, maximum likelihood estimation is performed based on the expected likelihood function of the hidden variable, and bag-of-word feature denoising training and label denoising training are included.

Benefits of technology

The convergence speed of the bag-of-word processing model is improved, and the tolerance and robustness of the model to feature and label noise are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114969327B_ABST
    Figure CN114969327B_ABST
Patent Text Reader

Abstract

The disclosed embodiments provide a training method and device for a bag-of-words processing model, and a bag-of-words processing method and device, wherein the bag-of-words processing model includes a posterior distribution model, wherein the posterior distribution model includes a first bag-of-words, a first label, a first latent variable, and a second latent variable; the first latent variable is a feature mask of the first bag-of-words, and the second latent variable is a feature mask of the first label; the bag-of-words processing model is trained based on an expectation-maximization (EM) framework, and the method includes: in an E step, based on an update rule of variational inference, optimizing the posterior distribution model by a fixed point iteration method, and calculating the expectation of the first latent variable and the second latent variable; in an M step, performing maximum likelihood estimation based on the expected likelihood function of the first latent variable and the second latent variable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of bag-of-words processing, and in particular to a training method and device for a bag-of-words processing model, and a bag-of-words processing method and device. Background Art

[0002] During the training process of the bag-of-words processing model, the model converges slowly. In addition, the presence of noisy variables and noisy labels will seriously affect the training process and generalization performance of the model. Summary of the Invention

[0003] In order to solve the above problems, the embodiments of the present disclosure provide a method and device for training a bag-of-words processing model, and a method and device for bag-of-words processing.

[0004] In a first aspect, a training method for a bag-of-words processing model is provided, wherein the bag-of-words processing model includes a posterior distribution model, and the posterior distribution model includes a first bag-of-words, a first label, a first latent variable, and a second latent variable; the first latent variable is a feature mask of the first bag-of-words, and the second latent variable is a feature mask of the first label; the bag-of-words processing model is trained based on an expectation-maximization (EM) framework, and the method includes: in the E step, based on the update rule of variational inference, optimizing the posterior distribution model by a fixed point iteration method, and calculating the expectation of the first latent variable and the second latent variable; in the M step, performing maximum likelihood estimation based on the expected likelihood function of the first latent variable and the second latent variable.

[0005] Optionally, the bag-of-words processing model also includes a predicted label model of the first bag-of-words, and the predicted label model is the probability distribution obeyed by the Hadamard product of the first bag-of-words and the first latent variable. Before step E, the method also includes: observing the sample data of the first bag-of-words and sampling the first latent variable multiple times; based on the sample data of the first bag-of-words and the multiple sampling data of the first latent variable, using the predicted label model to obtain multiple category label distributions of the first bag-of-words; performing Bayesian averaging on the multiple category label distributions of the first bag-of-words to obtain the predicted label distribution of the first bag-of-words.

[0006] Optionally, before step E, the method further includes: based on the sample data of the first word bag and the sampling data of the first latent variable, denoising the sample data of the first word bag by utilizing the Hadamard product of the first word bag and the first latent variable.

[0007] Optionally, both the first latent variable and the second latent variable obey Bernoulli distribution.

[0008] In a second aspect, a word bag processing method is provided, wherein the word bag processing model on which the word bag processing method is based includes a posterior distribution model, and the posterior distribution model includes a first word bag, a first label, a first latent variable, and a second latent variable; the first latent variable is a feature mask of the first word bag, and the second latent variable is a feature mask of the first label; the method includes: inputting the feature data of the first word bag to be processed into the word bag processing model to obtain the denoised word bag feature data.

[0009] Optionally, the bag-of-words processing model also includes a predicted label model of the first bag-of-words, and the predicted label model is the probability distribution obeyed by the Hadamard product of the first bag-of-words and the first latent variable. The method also includes: inputting the feature data of the first bag-of-words to be processed into the predicted label model to obtain the predicted label distribution of the first bag-of-words.

[0010] Optionally, both the first latent variable and the second latent variable obey Bernoulli distribution.

[0011] In a third aspect, a training device for a bag-of-words processing model is provided, wherein the bag-of-words processing model includes a posterior distribution model, and the posterior distribution model includes a first bag-of-words, a first label, a first latent variable, and a second latent variable; the first latent variable is a feature mask of the first bag-of-words, and the second latent variable is a feature mask of the first label; the bag-of-words processing model is trained based on an expectation-maximization (EM) framework, and the device includes: an optimization module for optimizing the posterior distribution model in an E-step, based on an update rule of variational inference, by adopting a fixed point iteration method, and calculating the expectation of the first latent variable and the second latent variable; an estimation module for performing maximum likelihood estimation in an M-step according to the expected likelihood function of the first latent variable and the second latent variable.

[0012] Optionally, the bag-of-words processing model also includes a predicted label model of the first bag-of-words, and the predicted label model is the probability distribution obeyed by the Hadamard product of the first bag-of-words and the first latent variable. The device also includes: an observation module, used to observe the sample data of the first bag-of-words before step E, and to sample the first latent variable multiple times; an acquisition module, used to obtain multiple category label distributions of the first bag-of-words based on the sample data of the first bag-of-words and multiple sampling data of the first latent variable using the predicted label model; an averaging module, used to perform Bayesian averaging on the multiple category label distributions of the first bag-of-words to obtain the predicted label distribution of the first bag-of-words.

[0013] Optionally, the device further includes: a denoising module, configured to denoise the sample data of the first word bag based on the sample data of the first word bag and the sampling data of the first latent variable by using the Hadamard product of the first word bag and the first latent variable before step E.

[0014] Optionally, both the first latent variable and the second latent variable obey Bernoulli distribution.

[0015] In a fourth aspect, a word bag processing device is provided, wherein the word bag processing model based on the word bag processing device includes a posterior distribution model, and the posterior distribution model includes a first word bag, a first label, a first latent variable and a second latent variable; the first latent variable is a feature mask of the first word bag, and the second latent variable is a feature mask of the first label; the device includes: an input module for inputting the feature data of the first word bag to be processed into the word bag processing model to obtain the denoised word bag feature data.

[0016] Optionally, the bag-of-words processing model also includes a predicted label model of the first bag-of-words, and the predicted label model is the probability distribution obeyed by the Hadamard product of the first bag-of-words and the first latent variable. The input module is used to: input the feature data of the first bag-of-words to be processed into the predicted label model to obtain the predicted label distribution of the first bag-of-words.

[0017] Optionally, both the first latent variable and the second latent variable obey Bernoulli distribution.

[0018] In a fifth aspect, a computer-readable storage medium is provided, on which executable code is stored. When the executable code is executed, the method described in the first aspect or the second aspect can be implemented.

[0019] In a sixth aspect, a computer program product is provided, comprising an executable code, which, when executed, can implement the method as described in the first aspect or the second aspect.

[0020] The embodiment of the present disclosure provides a training method for a bag-of-words processing model, which trains a bag-of-words processing model containing multiple latent variables based on the EM algorithm. In the E step, based on the update rule of variational inference, the posterior distribution model is optimized by fixed point iteration, which can improve the convergence speed of the bag-of-words processing model. In addition, the embodiment of the present disclosure provides a strict probabilistic form of bag-of-words feature noise and label noise. During the model training process, bag-of-words feature denoising training and label denoising training are included, which improves the model's tolerance and robustness to feature noise and label noise. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1A schematic diagram of the structure of a directed acyclic graph model provided in one embodiment of the present disclosure.

[0022] Figure 2 A schematic flowchart of a bag-of-words processing model training method provided in one embodiment of the present disclosure.

[0023] Figure 3 A schematic flowchart of a method for obtaining a predicted distribution according to an embodiment of the present disclosure.

[0024] Figure 4 A schematic flowchart of a bag-of-words processing model training method provided in another embodiment of the present disclosure.

[0025] Figure 5 A schematic flowchart of a bag-of-words processing method provided in one embodiment of the present disclosure.

[0026] Figure 6 A schematic diagram of the structure of a bag-of-words processing model training device provided in one embodiment of the present disclosure.

[0027] Figure 7 A schematic diagram of the structure of a bag-of-words processing device provided in one embodiment of the present disclosure.

[0028] Figure 8 A schematic structural diagram of a bag-of-words processing model training device provided in another embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments.

[0030] To facilitate understanding of the embodiments of the present application, the following is a brief introduction to the terms involved in the embodiments of the present application.

[0031] Hidden variables

[0032] A latent variable is a variable that cannot be directly observed in an experiment. As an example, when performing parameter estimation, we can use samples to find parameter estimates of the data distribution through the maximum likelihood method. However, if the samples come from multiple different distributions, and we do not know which distribution each sample comes from, traditional maximum likelihood estimation will not work. In this case, the distribution from which the sample comes becomes a latent variable, that is, the latent variable is the feature masking of the sample data. In other words, the latent variable in this example is used to mask the distribution information of the sample data.

[0033] EM algorithm

[0034] The EM algorithm (Expectation Maximization) is a classic algorithm in machine learning. It was proposed by Dempster in 1977. The EM algorithm is an iterative algorithm used to estimate the maximum likelihood or maximum a posteriori probability of probabilistic parameter models with latent variables. The EM algorithm is often used in data cluster analysis.

[0035] Variational Inference

[0036] Variational inference (VI) is a broad category of Bayesian approximate inference methods. Specifically, it is an iterative approximation algorithm that performs coordinate descent on the posterior distribution under the assumption of independence, cleverly transforming the posterior inference problem into an optimization problem for solution. In particular, when the posterior distribution model contains multiple hidden variables or is relatively complex, making it difficult to directly solve the posterior distribution, using variational inference to approximate the solution can achieve good convergence and scalability. As an example, the solution steps of variational inference can first construct an approximate distribution of the posterior probability distribution, and then use coordinate descent to continuously narrow the distance between the approximate distribution and the posterior probability distribution until convergence.

[0037] Currently, industry classification standards categorize companies into 20 primary industries, further subdivided into 96 secondary industries. A company's industry label is a crucial field, and with tens of millions of companies nationwide and many new businesses being created daily, quickly classifying them by industry is a crucial issue.

[0038] With the rapid development of machine learning, automated industry classification of enterprises using machine models has become widely used to address the tedious nature of traditional manual classification. An enterprise's industry classification is primarily derived from bag-of-words features, which are derived from the text describing the enterprise's business scope. However, these bag-of-words contain features that are irrelevant to the industry classification, meaning they are noisy. Furthermore, obtaining large amounts of labeled data is a key issue during model training. Manual labeling is time-consuming and labor-intensive, while some automated methods can quickly generate large amounts of samples but cannot guarantee accurate labeling and often contain some noise. Managing bag-of-words feature noise and label noise to provide more usable samples for downstream machine learning tasks is a pressing technical challenge.

[0039] In related technologies, Gibbs sampling is often used to approximate the posterior distribution during bag-of-words model training. However, this sampling method takes a long time to converge, sacrificing inference speed and performance, resulting in slow model convergence. Furthermore, the presence of noisy variables and labels can severely impact the model's training process and generalization performance, causing the model to enter a suboptimal steady state and reducing its applicability and robustness.

[0040] In order to solve the above problems, the embodiment of the present disclosure provides a training method for a bag-of-words processing model, which trains a bag-of-words processing model containing multiple latent variables based on the EM algorithm. Furthermore, in the solution of the embodiment of the present disclosure, in the E step, based on the update rule of variational inference, a fixed point iteration method is used to optimize the posterior distribution model, which can improve the convergence speed of the bag-of-words processing model. In addition, the embodiment of the present disclosure also provides a strict probabilistic form of bag-of-words feature noise and label noise. In the model training process, bag-of-words feature denoising training and label denoising training are included, which improves the model's tolerance and robustness to feature noise and label noise.

[0041] The bag-of-words processing model in the embodiment of the present disclosure may include a posterior distribution model. The posterior distribution model may be obtained based on sample data and a priori probability distribution, and is also called a posterior probability distribution or a joint posterior probability distribution. The posterior distribution model includes a first bag-of-words, a first label, a first latent variable, and a second latent variable. Among them, the feature data of the first bag-of-words may refer to a feature vector of length m (m is a positive integer); the label data of the first label may refer to the industry classification of the first bag-of-words. As an example, the feature data of the first bag-of-words may refer to the bag-of-words features composed of the business scope text of any enterprise in the industrial and commercial registration information, and the first label may refer to the industry classification of the first bag-of-words features in the enterprise's industrial and commercial registration information. The first latent variable may be a feature mask of the first bag-of-words, and the second latent variable may be a feature mask of the first label. It can be understood that the bag-of-words feature data of the first bag-of-words is discrete, and the label data of the first label is also discrete.

[0042] In some embodiments, the first latent variables in the present disclosure may all obey Bernoulli distribution. For example, the sample data of the first word bag can be represented by x ij To express, x ij Represents the jth word in the i-th word bag vector. Sample data vector x ij The latent variable (i.e. the first latent variable) can be expressed as z ij To express, z ij Indicates whether the jth word in the i-th bag-of-words vector is related to the industry classification. If z ij If it is 1, it means x ij is relevant, the jth word is not noise, otherwise it is considered as noise. ijFollows Bernoulli distribution:

[0043] p(z ii )=Bern(p)

[0044] Among them, i and j are both positive integers, and p in Bern(p) is a hyperparameter.

[0045] The word bag processing model also includes a prediction label model, for example, based on the sample data x of the first word bag i (can represent the i-th bag-of-words vector, the length of which can be m) and the first hidden quantity z i (vector z i Can be vector x i The feature mask of the prediction label model is constructed, and the prediction label y of the prediction label model is i ′ obeys the probability distribution with parameter θ:

[0046] p(y′ i |x i , z i )=p θ (y′ i |x i ⊙z i )

[0047] It can be understood that the above prediction label model is the sample data x of the first word bag i With the first hidden quantity z i The probability distribution obeyed by the Hadamard product of .

[0048] Furthermore, the predicted label y i The truth of ′ can be determined by the second latent variable c i To express, the second latent variable c i Can all obey Bernoulli distribution. The real label y i The probability distribution of can be expressed as a multinomial distribution (Mult):

[0049]

[0050] Among them, π is a hyperparameter.

[0051] It should be noted that c i Can be used to select the discriminant model. The discriminant model can be, for example, the multinomial distribution Mult(y′ i ) and Mult(π). As an example, when c i When it is 1, it means the label is clean. i When it is 0, it means the label is dirty.

[0052] In some embodiments, the posterior distribution model can be (y|x, c, z). It should be understood that the posterior distribution model can also be represented by the posterior distribution target p(z, c|x, y). Where x represents the sample data of the first bag of words, y represents the label data of the first label, c is the first latent variable, and z is the second latent variable.

[0053] In some embodiments, the present disclosure provides a directed acyclic graph for describing the above model. Figure 1 As shown in the figure, the directed acyclic graph model can intuitively describe the various parameters in the bag-of-words processing model (i.e., x ij , z ij , c i ,y i ′,y i ), where p and π are hyperparameters. In fact, the solution in the embodiment of the present disclosure is mainly based on the EM algorithm framework to train the graph model to obtain the latent variable parameters in the model.

[0054] It should be noted that the bag-of-words processing model in the present disclosure may include one or more latent variables, and the embodiments of the present disclosure do not impose specific limitations on this.

[0055] Combined with the following Figure 2 The training method of the bag-of-words processing model in the embodiment of the present disclosure is introduced. Figure 2 The method shown is a training method for a bag-of-words processing model based on an EM algorithm framework. The training method 200 includes steps S220 to S240.

[0056] Step S220: In step E, based on the update rule of variational inference, the posterior distribution model is optimized by fixed point iteration, and the expectations of the first latent variable and the second latent variable are calculated.

[0057] In some embodiments, in step E, a variational inference approach can be used to approximate the posterior distribution target p(z, c|x, y) of step E, that is, based on given observed sample data and sample label data (x, y), a posterior guess is made for the latent variable (z, c). For example, an approximate distribution of the posterior distribution target p(z, c|x, y) can be constructed based on an existing prior distribution, and then the coordinate descent method can be used to continuously narrow the distance between the approximate distribution and the posterior probability distribution until convergence. Transforming the problem of solving the posterior distribution into an optimization problem of the posterior distribution target reduces the difficulty of solving the posterior distribution target.

[0058] In other embodiments, in step E, the posterior distribution model can also be optimized by fixed point iteration based on the update rule of variational inference to improve the convergence speed of the model (the fixed point iteration method is essentially also to optimize the posterior distribution target, which will be combined later). Figure 4The method of fixed point iteration is illustrated by example. For details, please refer to Figure 4 The relevant description is not detailed here).

[0059] In some embodiments, before step E, based on the Hadamard product of the first bag of words and the first latent variable, the predicted label y can be obtained through the predicted label model. i ′. Thus, it can play the role of enhancing the data set. It should be understood that in order to improve the convergence speed of the word bag processing model, the acquisition of the predicted label y can be omitted. i As an example, during model training, the predicted label can be marginalized so that the predicted label y can be omitted in the E-step inference. i ′ process.

[0060] Step S240: In step M, maximum likelihood estimation is performed based on the expected likelihood function of the first latent variable and the second latent variable.

[0061] In some embodiments, the true label y i Obeying the probability distribution with parameter θ, the first latent variable and the second latent variable can be expressed by q t (z, c) represents, t is a positive integer, representing the number of iterations. t The expected likelihood function θ * For example, we can use the expected log-likelihood function:

[0062]

[0063] Among them, q * It means that in step Eq t The optimized variable representation is, is the expectation of the first latent variable and the second latent variable after E-step optimization.

[0064] In some embodiments, the expected log-likelihood function θ may be used in the M-step. * Under the framework of the EM algorithm, the bag-of-words processing model is trained by alternating two steps of EM until the expected likelihood function converges.

[0065] In some embodiments, based on the sample data of the first bag-of-words and the sample data of the first latent variable, the sample data of the first bag-of-words can be denoised using a Hadamard product of the sample data of the first bag-of-words and the sample data of the first latent variable, thereby improving the accuracy of the bag-of-words feature data.

[0066] In some embodiments, before step E, the sample data of the first word bag can be observed and the first latent variable can be sampled multiple times. Then, based on the sample data of the first word bag and the multiple sampled data of the first latent variable, the prediction label model can be used to obtain the multiple category label distributions of the first word bag. Then, the multiple category label distributions (also referred to as category distributions) of the first word bag are Bayesian averaged to obtain the predicted label distribution of the first word bag. Exemplarily, for example, the sample data of the first word bag can be Hadamard multiplied with the multiple sampled data of the first latent variable, and then the prediction label model can be used to obtain the multiple category label distributions of the first word bag.

[0067] The following combination Figure 3 The method for obtaining the predicted label distribution of the first bag of words is introduced in detail.

[0068] like Figure 3 As shown, before the E step, the first latent variable z can be i Perform multiple sampling (for example, three samplings) to obtain the feature vector x of the first word bag i Then, continue the forward calculation and use the posterior probability distribution p θ (y′ i |x i ⊙z i ), we can get the three category label distributions of the first word bag. Then, we perform Bayesian averaging on the three category label distributions of the first word bag to get x i The category prediction distribution (also called the predicted label distribution) can also be called the predicted label of the first bag of words. This can enhance the dataset. It should be understood that during the model training process, the predicted label can be marginalized, so that the predicted label y can be omitted in the E-step inference. i ′ process, improving the training convergence speed.

[0069] Furthermore, based on the feature vector x of the first bag of words i and the sampled data z of the first latent variable i , the feature data of the first word bag can be denoised. For example, the feature vector x of the first word bag can be i and the sampled data z of the first latent variable i Perform Hadamard product to denoise the feature data of the first word bag.

[0070] According to the above content, the training method of the bag-of-words processing model in the embodiment of the present disclosure is to train the bag-of-words processing model containing multiple latent variables based on the EM algorithm. In the E step, based on the update rule of variational inference, the posterior distribution model is optimized by fixed point iteration, which can improve the convergence speed of the bag-of-words processing model. In addition, for the bag-of-words feature data with discrete characteristics, a strict probabilistic form of bag-of-words feature noise and label noise is given in the embodiment of the present disclosure. In the model training process, bag-of-words feature denoising training and label denoising training are included, which improves the model's tolerance and robustness to feature noise and label noise.

[0071] For the convenience of description, we can construct the variable representation q of the first latent variable and the second latent variable t (z i , x i ), t is a positive integer, which represents the number of iterations of the first latent variable and the second latent variable. Figure 4 , a detailed description is given of the training method of the bag-of-words processing model in the embodiment of the present disclosure. Figure 4 The training method of the bag-of-words processing model in is still based on the EM algorithm framework.

[0072] In step E:

[0073] (1) The original data for model training may be obtained first. The original data may refer to the observed sample data of the first word bag and the observed label data of the first label.

[0074] (2) Initialize the parameters of the bag-of-words processing model. For example, you can initialize the latent variable q t , to get the initialized parameter q 0 (z i , x i ). It can be understood that this initialization step is to assign the hidden variable q t The initial value of .

[0075] (3) Construct a contraction mapping and apply it to the current latent variable q t On the other hand, fixed point iteration is used to calculate the hidden variable q by coordinate descent method. t Optimize to obtain the optimized latent variable parameter q * The update rule for each coordinate can be given by variational inference.

[0076] q t =Ψ[q t-1 θ * ]

[0077] It can be seen that each time q t The update of the global depends on the optimized latent variable q in the previous step t-1and the maximum likelihood estimate of the expected likelihood function of the latent variable θ * .

[0078] In the M step:

[0079] Get the hidden variable parameter q after E-step optimization * After that, in the latent variable parameter q * Under the expected log-likelihood function, the conventional stochastic gradient method is used to calculate the maximum likelihood estimate θ of the expected log-likelihood function * .

[0080]

[0081] By alternating and iteratively updating the EM two-step until the likelihood function converges, a bag-of-words processing model that meets the requirements can be obtained.

[0082] During the model training process, a strict probabilistic form of bag-of-words feature noise and label noise is given. The bag-of-words processing model trained by the above method has good robustness.

[0083] Based on the bag-of-words processing model trained in the previous article, the present disclosure also discloses a bag-of-words processing method. Figure 5 The bag-of-words processing method in the embodiment of the present invention is based on a bag-of-words processing model including a posterior distribution model, which includes a first bag-of-words, a first label, a first latent variable, and a second latent variable. The first latent variable is a feature mask of the first bag-of-words, and the second latent variable is a feature mask of the first label. It should be understood that the models involved in the bag-of-words processing method 500 are all trained models, that is, all latent variables and model parameters are known.

[0084] The bag-of-words processing method 500 includes step S520 .

[0085] In step S520 , the feature data of the first word bag to be processed is input into a word bag processing model to obtain denoised word bag feature data.

[0086] In some embodiments, the feature data of the first bag-of-words can be denoised, for example, by performing a Hadamard product of the feature vector data of the first bag-of-words with the vector data of the first latent variable to obtain the denoised feature data of the first bag-of-words. This can achieve a denoising effect on the bag-of-words feature data and improve the accuracy of the bag-of-words feature prediction label.

[0087] Optionally, the bag-of-words processing model also includes a predicted label model of the first bag-of-words, which is a probability distribution obeyed by the Hadamard product of the first bag-of-words and the first latent variable. The bag-of-words processing method 500 also includes: inputting the feature data of the first bag-of-words to be processed into the predicted label model, and then, the predicted label distribution of the first bag-of-words can be obtained.

[0088] Optionally, both the first latent variable and the second latent variable obey Bernoulli distribution.

[0089] It should be understood that the model in the bag-of-words processing method corresponds to the description of the bag-of-words model embodiment in the previous training method. Therefore, for parts not described in detail, please refer to the previous training method embodiment.

[0090] It can be understood that the model usage method and model training method in the embodiments of the present disclosure are not limited to the training of the bag-of-words processing model. For example, they can also be applied to discrete image noise processing models, discrete signal processing models, etc. The present disclosure does not impose specific restrictions on this.

[0091] Combined with the above Figures 1 to 5 , describes in detail the training method of the word bag processing model and the embodiment of the word bag processing method in the present disclosure, and Figures 6 to 8 , the device embodiment of the present disclosure is described in detail. It should be understood that the description of the device embodiment corresponds to the description of the method embodiment, so that parts not described in detail can refer to the previous method embodiment.

[0092] Figure 6 This is a schematic structural diagram of a training apparatus for a bag-of-words processing model provided in an embodiment of the present disclosure. The bag-of-words processing model includes a posterior distribution model, which includes a first bag-of-words, a first label, a first latent variable, and a second latent variable; the first latent variable is a feature mask of the first bag-of-words, and the second latent variable is a feature mask of the first label; the bag-of-words processing model is trained based on an expectation-maximization (EM) framework, and the apparatus 600 may include an optimization module 610 and an estimation module 620.

[0093] The optimization module 610 may be configured to optimize the posterior distribution model in step E based on an update rule of variational inference using a fixed point iteration method, and calculate the expectations of the first latent variable and the second latent variable.

[0094] The estimation module 620 may be configured to perform maximum likelihood estimation according to the expected likelihood function of the first latent variable and the second latent variable in the M-step.

[0095] Optionally, the bag-of-words processing model also includes a predicted label model of the first bag-of-words, and the predicted label model is the probability distribution obeyed by the Hadamard product of the first bag-of-words and the first latent variable. The device 600 also includes: an observation module 630 can be used to observe the sample data of the first bag-of-words before step E and perform multiple samplings on the first latent variable; an acquisition module 640 can be used to obtain multiple category label distributions of the first bag-of-words based on the sample data of the first bag-of-words and multiple sampling data of the first latent variable using the predicted label model; and an averaging module 650 can be used to perform Bayesian averaging on the multiple category label distributions of the first bag-of-words to obtain the predicted label distribution of the first bag-of-words.

[0096] Optionally, the device 600 also includes: a denoising module 660 can be used to denoise the sample data of the first word bag based on the sample data of the first word bag and the sampling data of the first latent variable before the E step, using the Hadamard product of the first word bag and the first latent variable.

[0097] Optionally, both the first latent variable and the second latent variable obey Bernoulli distribution.

[0098] Figure 7 : This is a schematic structural diagram of a bag-of-words processing apparatus provided in an embodiment of the present disclosure. The bag-of-words processing apparatus is based on a bag-of-words processing model including a posterior distribution model, which includes a first bag-of-words, a first label, a first latent variable, and a second latent variable; the first latent variable is a feature mask of the first bag-of-words, and the second latent variable is a feature mask of the first label; the apparatus 700 includes an input module 720.

[0099] The input module 720 may be configured to input the feature data of the first word bag to be processed into the word bag processing model to obtain the denoised word bag feature data.

[0100] In some embodiments, the feature data of the first bag-of-words can be denoised, for example, by performing a Hadamard product of the feature vector data of the first bag-of-words with the vector data of the first latent variable to obtain the denoised feature data of the first bag-of-words. This can achieve a denoising effect on the bag-of-words feature data and improve the accuracy of the bag-of-words feature prediction label.

[0101] Optionally, the bag-of-words processing model also includes a predicted label model of the first bag-of-words, and the predicted label model is the probability distribution obeyed by the Hadamard product of the first bag-of-words and the first latent variable. The input module 720 can be used to: input the feature data of the first bag-of-words to be processed into the predicted label model to obtain the predicted label distribution of the first bag-of-words.

[0102] Optionally, both the first latent variable and the second latent variable obey Bernoulli distribution.

[0103] Figure 8 It is a schematic structural diagram of the training device of the word bag processing model provided by the embodiment of the present disclosure. Figure 8 The dotted line in the figure indicates that the unit or module is optional. The device 800 can be used to implement the method described in the method embodiment above (i.e., any possible bag-of-words processing model training method and / or bag-of-words processing method described above). The device 800 can be a chip, a terminal, or a network device.

[0104] The device 800 may include one or more processors 810. The processor 810 may support the device 800 to implement the method described in the method embodiment above. The processor 810 may be a general-purpose processor or a special-purpose processor. For example, the processor may be a central processing unit (CPU). Alternatively, the processor may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0105] The apparatus 800 may further include one or more memories 820. The memories 820 store programs that can be executed by the processor 810, causing the processor 810 to perform the methods described in the above method embodiments. The memories 820 may be independent of the processor 810 or integrated into the processor 810.

[0106] The apparatus 800 may further include a transceiver 830. The processor 810 may communicate with other devices or chips via the transceiver 830. For example, the processor 810 may transmit and receive data with other devices or chips via the transceiver 830.

[0107] The embodiments of the present disclosure further provide a computer-readable storage medium having executable code stored thereon. When the executable code is executed, the methods described in the above-mentioned various method embodiments can be implemented.

[0108] The embodiments of the present disclosure further provide a computer program product having executable codes stored thereon. When the executable codes are executed, the methods described in the above-mentioned various method embodiments can be implemented.

[0109] The embodiments of the present disclosure further provide a computer program, which, when executed, can implement the methods described in the above-mentioned various method embodiments.

[0110] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any other combination. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0111] Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with the embodiments of the present disclosure can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0112] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0113] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0114] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0115] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A training method for a bag-of-words processing model, the bag-of-words processing model comprising a posterior distribution model, the posterior distribution model comprising a first bag-of-words, a first label, a first latent variable, and a second latent variable; the first latent variable is a feature mask of the first bag-of-words, and the second latent variable is a feature mask of the first label; The bag-of-words processing model is trained based on an expectation-maximization (EM) framework, and the method includes: In step E, based on the update rule of variational inference, the posterior distribution model is optimized by fixed point iteration, and the expectations of the first latent variable and the second latent variable are calculated; In the M step, maximum likelihood estimation is performed based on the expected likelihood function of the first latent variable and the second latent variable.

2. According to the method of claim 1, the word bag processing model further includes a prediction label model of the first word bag, the prediction label model is a probability distribution obeyed by the Hadamard product of the first word bag and the first latent variable, and before the E step, the method further includes: Observe sample data of the first word bag, and sample the first latent variable multiple times; Based on the sample data of the first word bag and the multiple sampling data of the first latent variable, using the prediction label model to obtain multiple category label distributions of the first word bag; Bayesian averaging is performed on multiple category label distributions of the first word bag to obtain a predicted label distribution of the first word bag.

3. The method according to claim 1, before step E, the method further comprises: Based on the sample data of the first word bag and the sampling data of the first latent variable, the sample data of the first word bag is denoised by utilizing the Hadamard product of the first word bag and the first latent variable.

4. According to the method of claim 1, both the first latent variable and the second latent variable obey Bernoulli distribution.

5. A bag-of-words processing method, wherein the bag-of-words processing method is based on a bag-of-words processing model trained by the training method of the bag-of-words processing model according to any one of claims 1 to 4, and the bag-of-words processing model comprises a posterior distribution model, wherein the posterior distribution model comprises a first bag-of-words, a first label, a first latent variable, and a second latent variable; the first latent variable is a feature mask of the first bag-of-words, and the second latent variable is a feature mask of the first label; The method comprises: The feature data of the first word bag to be processed is input into the word bag processing model to obtain the denoised word bag feature data.

6. The method according to claim 5, wherein the bag-of-words processing model further comprises a prediction label model of the first bag-of-words, wherein the prediction label model is a probability distribution obeyed by the Hadamard product of the first bag-of-words and the first latent variable, and the method further comprises: The feature data of the first word bag to be processed is input into the predicted label model to obtain the predicted label distribution of the first word bag.

7. According to the method of claim 5, both the first latent variable and the second latent variable obey Bernoulli distribution.

8. A training device for a bag-of-words processing model, the bag-of-words processing model comprising a posterior distribution model, the posterior distribution model comprising a first bag-of-words, a first label, a first latent variable, and a second latent variable; the first latent variable is a feature mask of the first bag-of-words, and the second latent variable is a feature mask of the first label; The bag-of-words processing model is trained based on an expectation-maximization (EM) framework, and the device comprises: An optimization module, used for optimizing the posterior distribution model in a fixed point iteration manner based on an update rule of variational inference in step E, and calculating the expectations of the first latent variable and the second latent variable; An estimation module is used to perform maximum likelihood estimation according to the expected likelihood function of the first latent variable and the second latent variable in the M step.

9. The device according to claim 8, wherein the bag-of-words processing model further comprises a prediction label model of the first bag-of-words, wherein the prediction label model is a probability distribution obeyed by the Hadamard product of the first bag-of-words and the first latent variable, and the device further comprises: An observation module, used for observing the sample data of the first word bag and sampling the first latent variable multiple times before the E step; An acquisition module, configured to acquire a plurality of category label distributions of the first word bag using the prediction label model based on the sample data of the first word bag and a plurality of sampling data of the first latent variable; An averaging module is used to perform Bayesian averaging on multiple category label distributions of the first word bag to obtain a predicted label distribution of the first word bag.

10. The device according to claim 9, further comprising: A denoising module is used to denoise the sample data of the first word bag by using the Hadamard product of the first word bag and the first latent variable based on the sample data of the first word bag and the sample data of the first latent variable before the E step.

11. The apparatus according to claim 8, wherein the first latent variable and the second latent variable both obey Bernoulli distribution.

12. A bag-of-words processing device, wherein the bag-of-words processing device is based on a bag-of-words processing model trained by the training method of the bag-of-words processing model according to any one of claims 1 to 4, and the bag-of-words processing model comprises a posterior distribution model, wherein the posterior distribution model comprises a first bag-of-words, a first label, a first latent variable, and a second latent variable; the first latent variable is a feature mask of the first bag-of-words, and the second latent variable is a feature mask of the first label; The device comprises: An input module is used to input the feature data of the first word bag to be processed into the word bag processing model to obtain the denoised word bag feature data.

13. The device according to claim 12, wherein the bag-of-words processing model further comprises a prediction label model of the first bag-of-words, the prediction label model being a probability distribution obeyed by a Hadamard product of the first bag-of-words and the first latent variable, and the input module is used to: The feature data of the first word bag to be processed is input into the predicted label model to obtain the predicted label distribution of the first word bag.

14. The apparatus according to claim 12, wherein the first latent variable and the second latent variable both obey Bernoulli distribution.

Citation Information

Patent Citations

  • A method for caption annotation of an image based on an extended sLDA model

    CN108984726A

  • An information processing method and device and an information detection method and device

    CN109685087A