A niche preference learning method for generating content based on user text
By employing a niche preference learning method based on user-generated text content, and utilizing hierarchical Bayesian models and Gibbs sampling, the challenge of niche preference identification was solved. This enabled accurate analysis of niche markets and identification of target users, improving the accuracy and data credibility of SMEs entering niche markets.
Patent Information
- Application Number
- CN202310127387.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2043-02-17
AI Technical Summary
Existing research has not been successful in identifying and modeling niche preferences. It is difficult to accurately identify the meaning of niche preferences, and the data collection process is arduous, lacking, or unreliable, which affects the accuracy of SMEs entering niche markets.
We employ a niche preference learning method based on user-generated text content. Through data preprocessing and the establishment of a hierarchical Bayesian model, we use Gibbs sampling to learn the model parameters, identify the meaning of user niche preferences, and find target users.
It effectively distinguishes between mass preferences and niche preferences, identifies the specific meaning of niche preferences, provides opportunities for SMEs to enter suitable niche markets, and improves market accuracy and data credibility.
Smart Images

Figure CN116340498B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of information retrieval, and particularly relates to a method for learning small-group preferences based on user text generation content. BACKGROUND
[0002] The development of online shopping websites provides a convenient and fast channel for small and medium-sized enterprises to sell products and services. Since entering the large-scale mature market must face huge and powerful competitors and monopolists, small markets are suitable places for small and medium-sized enterprises to obtain small market shares and expand market shares in the future. Small and medium-sized enterprises choose small markets not only because of the advantages of less competition and relatively cheap marketing costs, but also because of the relatively high investment return and the opportunity for long-term success.
[0003] Previous studies on small markets have predetermined small markets or small products, usually based on experience and prior knowledge, and using classification methods to select suitable small markets is a common means of research. Subsequently, according to the pre-determined small market and small product, the analysis and identification of preferences are the usual operation. However, small markets are the result of small products, and small products are produced according to the influence of small-group preferences and ultimately, helping small and medium-sized enterprises to better enter suitable small markets and produce products, the order of research should be to analyze small-group preferences first, and then discover the market.
[0004] Although the analysis of small-group preferences can enhance the understanding of users and small and medium-sized enterprises that try to easily enter small markets, existing research is not successful in determining and modeling small-group preferences. The identification and modeling of small-group preferences include two main goals, namely obtaining the preference distribution of users and understanding its specific meaning. Some methods can accurately learn the small-group preference distribution, but the meaning of small-group preferences as the central point of discovering potential small markets cannot be identified. Some scholars choose and collect product features, such as price range, production quality, and demographic data, to understand the meaning of small-group preferences. Although the data of users can be used to understand and identify small-group preferences, the process of collecting data is arduous, and the data is often lacking or unreliable. SUMMARY
[0005] In view of the deficiencies of the prior art, the purpose of the present application is to provide a method for learning small-group preferences based on user text generation content to solve the problems raised in the background art.
[0006] The purpose of the present application can be achieved by the following technical solutions:
[0007] A method for learning small-group preferences based on user text generation content, comprising the following steps:
[0008] Performing data preprocessing operations on the obtained user text generation content;
[0009] The data obtained by the preprocessing is used to establish a hierarchical Bayesian model to obtain a joint distribution model;
[0010] The model parameters are learned by Gibbs sampling method to obtain the mass preference distribution and the minority preference distribution formula;
[0011] The learned model parameters are used to analyze the meaning of the user's minority preference for generating content based on the user's text;
[0012] The target user under the minority preference is found by using the user's minority preference distribution.
[0013] Preferably, the data preprocessing operation first removes the non-text part of the obtained document, then performs a word segmentation operation on the document, and finally performs cleaning work on the segmented document to obtain U documents after data preprocessing.
[0014] Preferably, the joint distribution obtained in the hierarchical Bayesian model is as follows:
[0015]
[0016] In the formula, w represents a word in the document; Z * ,Z * represents the preference of each word, the front represents the mass preference of the word, and the back represents the minority preference; y is a binary variable, indicating whether the word generation process is affected by the mass preference or the minority preference; α * ,α * ,β * ,β * ,γ0,γ1 are hyperparameters of the prior distribution; first represents the joint distribution of the binary variable y; second represents the mass preference distribution of the user u; third represents the distribution of the word under the mass preference z * ; fourth represents the minority preference distribution of the user u; and fifth represents the distribution of the word under the minority preference z * .
[0017] Preferably, the step 3 needs to learn and solve first, second, third, fourth, and fifth in the joint distribution;
[0018] The solving formula of the first is as follows:
[0019]
[0020]
[0021] In the formula, n u (y)represents the number of times the word w appears in the document u; θ
[0022] The solution for the second is as follows:
[0023]
[0024]
[0025] where n *,u(w) (m) represents the number of times the word w appears in the document u; θ * represents the document-popularity distribution;
[0026] The solution for the third is as follows:
[0027]
[0028]
[0029] where n *,m (v,~) represents the number of times the word w appears in the document u; θ * represents the document-popularity distribution;
[0030] The solution for the fourth is as follows:
[0031]
[0032]
[0033] where n represents the number of times the word w appears in the document u; θ * represents the document-popularity distribution;
[0034] The solution for the fifth is as follows:
[0035]
[0036]
[0037] where n represents the number of times the word w appears in the document u; θ * represents the document-popularity distribution;
[0038] Preferably, the popularity distribution is as follows:
[0039]
[0040] The niche preference distribution formula is as follows:
[0041]
[0042] Preferably, the parameter results of the step 3 model parameters are as follows:
[0043]
[0044]
[0045]
[0046]
[0047]
[0048] Preferably, the meaning identification process of the user's niche preference in the step 4 is as follows:
[0049] The user's niche preference includes two parts, i.e. a document-niche preference distribution and a niche preference-word distribution The words in each niche preference word distribution are sorted according to the probability size, i.e. The probabilities under different words are selected to analyze the first ten words, so as to identify the specific meaning of the niche preference.
[0050] Preferably, the process of finding target users under the niche preference in the step 5 is as follows:
[0051] The user's niche preference distribution is used to find target users under the niche preference, and the niche preference distribution of each user is the document niche preference distribution, as follows:
[0052] d i =[p(z1|d i ),…,p(z k |d i )]
[0053] Randomly selected users are analyzed under different niche preferences, so as to find target users under different niche preferences.
[0054] The present application has the following advantages:
[0055] 1. The method distinguishes between mass preferences and niche preferences from the perspective of user preferences, uses the good interpretability of the hierarchical Bayesian method to identify the specific meaning of user niche preferences, provides small and medium-sized enterprises with opportunities to enter suitable niche markets, and the distribution of each user's niche preferences helps enterprises find target users in related niche markets. BRIEF DESCRIPTION OF DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0057] Figure 1 is a flowchart of the method of the present application;
[0058] Figure 2 is a confusion index diagram for determining the number of niche preferences on the Douban dataset according to the present application;
[0059] Figure 3 is a table of the first ten words and the identified meanings under the niche preference found on the Douban dataset according to the present application;
[0060] Figure 4 is a probability table learned by the present application on the Douban dataset for different users under different niche preferences. DETAILED DESCRIPTION
[0061] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0062] The present application proposes a niche preference learning method based on user text generated content, comprising the following steps:
[0063] Step 1, data preprocessing operation is performed on the user text generated content obtained;
[0064] Step 2, a hierarchical Bayesian model is established for the data obtained by preprocessing in step 1;
[0065] Step 3, the model parameters of step 2 are learned by Gibbs sampling method;
[0066] Step 4, the meaning of user niche preference based on user text generated content is analyzed by using the model parameters learned in step 3;
[0067] Step 5, find target users under the niche preference by using the user niche preference distribution.
[0068] Wherein, step 1 is as follows:
[0069] Obtain the text generation content of the user on the social platform, synthesize all the texts published by each user into a document, and finally obtain U user documents. Perform data preprocessing operations on the U documents. First, remove the non-text part of the obtained document, such as web addresses; then perform word segmentation on the document. Chinese documents can use tools such as Jieba for word segmentation, and English documents with special formats can not be segmented; finally, clean the segmented document to remove useless labels, special symbols, stop words, etc., and obtain the U documents after data preprocessing.
[0070] Step 2 is as follows:
[0071] Establish a hierarchical Bayesian model for the U documents obtained in step 1. Before establishing the model, define the related symbols that need to be used later, as shown in the following table.
[0072] Table 1: Symbols and descriptions used in patents
[0073]
[0074] In the construction of the hierarchical Bayesian model, it is assumed that there are two preferences in all documents, M popular preferences and K niche preferences; wherein, the popular preference and the niche preference have a multivariate distribution ψ * and ψ * Each document has two multivariate distributions, document-popular preference distribution θ * and document-niche preference distribution θ * ; whether a word is generated by a niche preference is determined by a Bernoulli distribution π based on an indicator y; the document preference distributions θ * and θ * have Dirichlet prior parameters α * and α * , the preference word distributions ψ * and ψ * have Dirichlet prior parameters β * and β * , and the Bernoulli distribution π of the word has beta prior parameters γ0 and γ1.
[0075] Based on the above assumptions, the hierarchical Bayesian model is constructed using the following generation process:
[0076] For each popular preference m∈[1,M] and each niche preference k∈[1,K]
[0077] Sampling popular preference distribution through word ψ m ~Dirichlet(β) * )
[0078] Sampling of niche preferences distribution through word ψ k ~Dirichlet(β) * )
[0079] For each document u∈[1,U]
[0080] Sampling Documents - Popular Preference Distribution θ * ~Dirichlet(α * )
[0081] Sampling Documents - Niche Preference Distribution θ * ~Dirichlet(α * )
[0082] The Bernoulli distribution parameter π ~ Beta(γ) of the sampled words
[0083] For each word w∈[1,V]
[0084] Sample a binary indicator y~Bernoulli(π).
[0085] If y = 1, then
[0086] Sampling yields popular word preferences z * ~Multinomial(θ) * )
[0087] sampling
[0088] If y = 0, then
[0089] Niche preferences for sampled words z * ~Multinomial(θ) * )
[0090] sampling
[0091] The generation process can be understood as follows: the generation of word w in the document first uses a Bernoulli distribution π to generate a binary indicator parameter y to determine which preference generates the word; if y = 1, then the word generation utilizes the user's mass preference, using the mass preference distribution θ. * Generate preference z * Then, the word w is generated using a multivariate distribution; if y = 0, the word w is generated by a niche preference distribution. Generate, z * It is a distribution of niche preferences θ *generated, resulting in a final joint distribution as follows:
[0092]
[0093] where first represents the joint distribution of binary variable y, second represents the mass preference distribution of user u, third represents the mass preference z * distribution of words, fourth represents the niche preference distribution of user u, and fifth represents the niche preference z * distribution of words.
[0094] Step 3 is detailed as follows:
[0095] The model parameters of Step 2 are learned by the Gibbs sampling method, which mainly uses the Gibbs sampling method to learn and solve first, second, third, fourth, and fifth in the joint distribution formula (1). The solution formula of first is shown in formula (2) as follows:
[0096]
[0097]
[0098] where n u (y) represents the number of times the word w appears in the mass preference m in the document u; the solution formula of third is shown in formula (4) as follows:
[0099]
[0100]
[0101] where n *,u(w) (m) represents the number of times the word w appears in the mass preference m in the document u; the solution formula of third is shown in formula (4) as follows:
[0102]
[0103]
[0104] where n *,m (v,~) represents the number of times the word w appears in the mass preference m in the document u; the solution formula of third is shown in formula (4) as follows:
[0105]
[0106]
[0107] where n represents the number of occurrences of word w in the k-th niche preference of user u; the solution formula of Fifth is shown in equation (6):
[0108]
[0109]
[0110] wherein, represents the number of occurrences of word v in the k-th niche preference, i represents the i-th word, and -i represents removing the i-th word from the document and the preference; the conditional distribution p(z *,i = m, y i = 1 | w, t, z *,-i , y -i , a * , b * , g) can be calculated by the above equations (2)-(6); for the mass preference, there is the following equation (7):
[0111]
[0112] For the niche preference distribution can be calculated as equation (8):
[0113]
[0114] Step 4 is as follows:
[0115] Using the model parameters learned in step 3, the parameter results of equations (9)-(13) can be obtained.
[0116]
[0117]
[0118]
[0119]
[0120]
[0121] The data used in this experiment is the user comment data on Douban, which analyzes the niche preferences of users based on user text generated content. The data collected from June 12, 2005 to September 7, 2019 contains 4707596 comments, and after data preprocessing, there are 31044125 valid comments, 17296360 words, and 34271 words. The average length of the comments is 5.57 words. In the experiment, b * = b * = 0.01, a * = 10-7 ,α * = 50 / k, y0= y1= 1, the iteration number is 1000, the number of topics is set in advance in the patent, the number of popular preferences is set to 1, the number of less popular preferences is learned by using the perplexity index, and the perplexity learning formula is shown in formula (14):
[0122]
[0123] In the formula, log(p(W u )) is the log function of the word probability, N u represents the number of test sets; in the experiment, 80% of the data is set as the training set, and 20% of the data is set as the test set; the smaller the perplexity index value is, the better the effect is, and finally the result of Figure 1 is obtained; according to the result of Figure 1 , the number of less popular preferences is selected as 60.
[0124] The less popular preferences of the user include two parts, the document-less popular preference distribution and the less popular preference-word distribution The words in each less popular preference word distribution are sorted according to the probability size, that is the probability under different words, the top ten words are selected for analysis, so as to identify the specific meaning of the less popular preferences, and several less popular preferences are selected for display and identification of the meaning, as shown in Figure 2 .
[0125] Step 5 is as follows:
[0126] The user less popular preference distribution is used to find target users under the less popular preferences, and the less popular preference distribution of each user is the document less popular preference distribution, as shown in formula (15); random users are selected, and their probabilities under different less popular preferences are analyzed, as shown in Figure 3 , so as to find target users under different less popular preferences. The higher the probability under the less popular preference is, the more target users of the less popular market discovered by the less popular preference.
[0127] d i = [p(z1|d i ), …, p(z k |d i )] (15)
[0128] In the description of the specification, reference to "one embodiment", "an example", "certain examples" etc. means that a particular feature, structure, material or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the application. The appearances of the phrases "in one embodiment" or "an example", "in certain embodiments" or "certain examples" in various places in the specification are not necessarily all referring to the same embodiment or example. Furthermore, the particular features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0129] Those skilled in the art will appreciate that embodiments of the application can be devised for a method, a system, or a computer program product. Accordingly, the application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the application can take the form of a computer program product on one or more computer readable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code thereon.
[0130] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks or in conjunction with the flowcharts described above. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks or in conjunction with the flowcharts described above.
[0131] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks or in conjunction with the flowcharts described above. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks or in conjunction with the flowcharts described above.
[0132] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks or in conjunction with the flowcharts described above. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart block or blocks or in conjunction with the flowcharts described above.
[0133] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application but not to limit it. Although the present application has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the specific embodiments of the present application can be modified or equivalently replaced without departing from the spirit and scope of the present application, and any modification or equivalent replacement should be covered in the protection scope of the claims of the present application.
Claims
1. A method for learning niche preferences based on user text generation of content, characterized in that, Comprising the following steps: Step 1, data preprocessing operation is generated for the user text acquisition content; Step 2, the data obtained by step 1 preprocessing is established a hierarchical Bayesian model, and the joint distribution model is obtained; Step 3, the model parameters of step 2 are learned by Gibbs sampling method, and the mass preference distribution and the minority preference distribution formula are obtained; Step 4, the meaning of the user's minority preference based on the user text generation content is analyzed by using the model parameters learned in step 3; Step 5, the target user under the minority preference is found by using the user minority preference distribution. The joint distribution obtained in the hierarchical Bayesian model is as follows: where w represents a word in a document; represents the preference that each word belongs to, the former represents that the word belongs to the popular preference, and the latter represents that the word belongs to the unpopular preference; y is a binary variable indicating whether the word generation process is influenced by the mass preference or the niche preference; is a hyperparameter for the prior distribution; first represents the joint distribution of the binary variable y; second represents the mass preference distribution of the user u; third represents the mass preference distribution of the word; fourth represents the niche preference distribution of the user u; fifth represents the niche preference distribution of the word under the niche preference; The solving formula of the Fourth is as follows: wherein, represents the number of occurrences of the word w in the niche preference k in the user u; represents the document-niche preference distribution; The solving formula of the Fifth is as follows: wherein, represents the number of occurrences of the word v in the niche preference, represents the i-th word, represents the removal of the i-th word from the document and the preference; represents the niche preference-word distribution. 2.The method of claim 1, wherein, The data preprocessing operation firstly removes the non-text part of the obtained document, then performs word segmentation operation on the document, and finally cleans the segmented document to obtain U documents after data preprocessing. 3.The method of claim 1, wherein, The solving formula of the first is as follows: wherein denotes the number of occurrences of the exponent y in the user u; is a Bernoulli distribution for each word; The solving formula of the second is as follows: wherein, represents the number of occurrences of a word w in the popular preference m in the document u; represents the document-popular preference distribution; The solving formula of the Third is as follows: wherein, represents the number of occurrences of the word v in the mass preference m; represents the mass preference-word distribution. 4.The method of claim 3, wherein, The mass preference distribution formula is as follows: The minority preference distribution formula is as follows: 。 5.The method of claim 4, wherein, The parameter results of the model parameters of step 3 are as follows: 。 6.The method of claim 5, wherein, The meaning recognition process of the user's minority preference in step 4 is as follows: The user's niche preferences include two parts: document-niche preference distribution and niche preference-word distribution The words in each niche preference-word distribution are sorted according to the probability size, that is The probabilities under different words are selected, and the top ten words are analyzed to identify the specific meaning of the niche preference.
7. The method of claim 6, wherein the method further comprises: The process of finding the target user under the minority preference in step 5 is as follows: Utilizing user niche preference distribution Finding target users under niche preference, each user's niche preference distribution is the document's niche preference distribution, as follows: Randomly select users, analyze their probability under different minority preferences, and find the target user under different minority preferences.
8. A system for performing the method of learning the niche preference based on the user text generation content according to any one of claims 1-7, characterized in that, Comprise: Data acquisition preprocessing module, used for acquiring document information, and screening and cleaning the acquired document information to obtain preprocessed data; Data processing module, used for distinguishing the mass preference distribution and the minority preference distribution of the data preprocessed by the data acquisition preprocessing module; Analysis module, used for analyzing the meaning of the user's minority preference based on the user text generation content; Finding module, used for finding the target user under the minority preference by using the user minority preference distribution.
9. A user text generation content based minority preference learning controller, storing a program for running the user text generation content based minority preference learning method of any one of claims 1-7.
Citation Information
Patent Citations
Personalized event recommendation method and system fusing theme matching and bidirectional preferences
CN111428127A
Generating a Targeted Summary of Textual Content Tuned to a Target Audience Vocabulary
US20190155877A1