Knowledge and Interest Personalized Profiling System for Internet Users
Through LDA improvement models and user interest and knowledge-level tracing systems, the problem that Internet users find it difficult to find accurate information in massive information is solved, and the accuracy of accurate personalized searches and e-commerce recommendations is improved.
Patent Information
- Application Number
- CN202011641277.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-31
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-12-31
AI Technical Summary
The existing technology cannot effectively analyze the personal knowledge system of Internet users, which makes it difficult for users to quickly find accurate information to meet their needs in massive information. The search results do not match the user's intentions, affecting the accuracy of user experience and e-commerce recommendations.
The Internet user knowledge and interest tracing system based on LDA improvement model is adopted, and the user interest and knowledge hierarchy model is constructed by analyzing the user's network usage history and Weibo data, combining keyword models and directed connectivity diagrams to accurately characterize user characteristics.
It realizes a comprehensive analysis of user behavior characteristics, interest distribution and knowledge fields, provides accurate and personalized search services, improves the fit and user experience of search results, and enhances the accuracy of e-commerce recommendations.
Smart Images

Figure CN112733021B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a system for depicting the knowledge and interests of users, and particularly to a personalized system for depicting the knowledge and interests of Internet users, belonging to the technical field of Internet user depiction. Background Art
[0002] In recent years, with the rapid development of the Internet, especially mobile Internet, including wireless WIFI, 4G networks provided by telecom service providers, and particularly the current mobile 5G networks strongly promoted by major telecom service providers, all indicate the great trend of the future high-speed development of user-centered mobile social networks and e-commerce. Therefore, the analysis of Internet user profiling has begun to receive key attention. Internet user personalized profiling calculation is to comprehensively analyze the user behavior characteristics, interest distribution, and the distribution of the user's existing knowledge fields based on the Internet user's network usage history. The specific difficulties lie in the representation of behavior, the quantification of interests, the measurement of knowledge fields, and the research and analysis of the combination points among the three.
[0003] Relying on the popularization of the global mobile information network and the accelerating progress of the informatization process, the amount of user data generated by mobile Internet sites, portal websites, company websites, blogs, microblogs, forums and other online media, as well as various mobile APPs, is huge and very unevenly distributed. However, for undifferentiated search, major search engines have provided solutions to the problem of quickly searching for relevant results that match the query. However, with the continuous development of mobile Internet applications and big data technology, and the general trend of mobile terminals leading to a rapid growth in the number of mobile Internet users, coupled with the fact that network users have put forward higher requirements for Internet knowledge acquisition, future mobile search will surely develop in the direction of intelligent and knowledge-based search engines, aiming to provide an intelligent search function that can provide a more accurate event summary of event elements that meets the requirements of user personalized search and understands user logic.
[0004] As the mobile Internet media that generates the most concentrated user data, mobile social network APPs, as an application service, can help people quickly and conveniently initiate social network activities. The forms of mobile social networks are diverse, and the diverse data generated during this long process can precisely provide support for analyzing and studying user personalized characteristics, such as user interest distribution, user behavior analysis, and the user's existing knowledge, in terms of a rich type of analysis data; especially in the case of the continuous development of mobile terminals and the explosive growth in the number of mobile terminals, various Internet companies have provided a large number of mobile APPs, generating a large amount of user personal text data to provide sufficient data volume support for the analysis and research of user personalized characteristics.
[0005] As a new type of self-media, Weibo has attracted a great deal of analysis and research on Internet users during its rise and booming development. However, little attention has been paid to the analysis of the existing knowledge systems of individual users. Nevertheless, this user characteristic is a crucial factor in aspects such as reflecting the true search intentions of Internet users during the process of using search engines, the selection of professional products in online shopping, and the quality of the generated text information. With the rapid development of the Internet, e-commerce has also developed at an extremely fast pace, and the number of e-commerce websites of various categories has been increasing day by day. The competition among e-commerce companies is becoming increasingly fierce. How to convert potential users in front of mobile phone screens into actual consumers and how to recommend more products that they are interested in to registered users of the website and stimulate consumption have also become challenges faced by e-commerce companies, and this has also become a key technology for gaining a foothold in this highly competitive industry.
[0006] With the advent of a more abundant and convenient mobile Internet era, people's demand for obtaining information is getting stronger and the standards are getting higher. User personalized services can well coordinate the contradiction between the explosive growth of information and people's need to obtain specific content in the fields they truly care about.
[0007] In the field of e-commerce, the first thing to do to achieve the above goals is to deeply analyze the characteristics of Internet users, such as their interests and knowledge levels. Because this is the prerequisite for correctly recommending potential purchase products to users. If this cannot be achieved, users will only be submerged in a vast ocean of products that are dazzling or even completely do not meet their needs. Continuing like this will only make users spend more time screening and searching again and again, wasting a lot of unnecessary energy to find products that meet their own needs. Such an experience is quite bad for users and also does not create the value it should for enterprises, resulting in a huge loss for all parties.
[0008] In the field of search engines, in order to screen out the most needed search results for users from the overly redundant and complex information, it is also necessary to be based on aspects such as users' interests and knowledge. Integrating this information can overall depict a real and specific user, and the returned search results will also be more in line with the user's intentions based on the understanding of the user. In this way, users no longer need to spend a lot of unnecessary time filtering the still messy search results twice or even three times to find the information they truly need, which not only saves users' time but also saves their energy, and also plays a huge role in improving and promoting the user experience.
[0009] Based on the current situation and deficiencies of existing technologies, the problems to be solved by the present invention are manifested as follows:
[0010] First, with the rapid development of the mobile Internet and the explosive growth of data volume, the vast amount of information has brought great difficulties to rapid information search. Users' requirements for information acquisition are also getting higher and higher. More and more users are not satisfied with simple information search services and need to quickly and efficiently find accurate information that meets their needs. This has brought great challenges to data services. The existing technologies do not pay enough attention to the existing knowledge distribution of users. The problem that the vast amount of information brings great difficulties to rapid information search cannot meet users' high requirements for information acquisition, cannot provide accurate personalized information search services, and users cannot quickly and efficiently find accurate information that meets their needs;
[0011] Second, the existing technologies rarely pay attention to the analysis and research of users' personal existing knowledge systems. However, this user characteristic is a very crucial factor in aspects such as reflecting the true search intentions of network users during the process of using search engines, choosing professional products in online shopping, and the quality of generated text information. The competition among current e-commerce platforms is becoming increasingly fierce. How to convert potential users in front of mobile phone screens into actual consumers, and how to recommend more products that they are interested in to website registered users and stimulate consumption have also become challenges faced by e-commerce enterprises. This has also become the key technology for gaining a foothold in this highly competitive industry. The existing technologies cannot deeply analyze the characteristics of Internet users. Users are submerged in a flood of dazzling products that do not even meet their needs at all. Users spend more time screening and searching again and again, wasting a lot of unnecessary energy to find products that meet their needs. Such an experience is quite bad for users and does not create the value it should for enterprises, resulting in a huge loss for all parties;
[0012] Third, in the field of search engines, in order to deal with overly redundant and complex information, the existing technologies cannot screen out the search results that users currently need the most, cannot overall depict a real and specific user, and the returned search results do not match users' intentions. Users need to spend a lot of unnecessary time to filter the still messy search results for the second or even third time to find the information they really need, wasting both users' time and energy. For users' usage experience, it is also very bad. Future mobile search needs to develop in the direction of intelligent and knowledge-based search engines, providing intelligent search functions that are more in line with users' personalized search requirements and can understand the event element type of users' logic and provide accurate event summaries;
[0013] Fourth, the problem that a vast amount of information poses great difficulties for rapid information search. The existing technologies cannot meet the increasingly high requirements of users for information acquisition, provide precise personalized information search services, and enable users to find accurate information that meets their needs more quickly and efficiently. The contradiction between the explosively growing information and the content in specific fields that people truly care about is extremely prominent. The existing technologies cannot comprehensively analyze users' behavioral characteristics based on their network usage history, nor can they comprehensively analyze users' behavioral characteristics, interest distributions, and the distributions of their personal existing knowledge fields through personalized computing for Internet users. They cannot optimize and guide the reflection of users' true search intentions during the use of search engines, the selection of professional products in online shopping, and the quality of the generated text information. Summary of the Invention
[0014] Aiming at the deficiencies of the existing technologies, the personalized knowledge and interest profiling system for Internet users provided by the present invention solves the problem that a vast amount of information poses great difficulties for rapid information search, meets the increasingly high requirements of users for information acquisition, provides precise personalized information search services, enables users to find accurate information that meets their needs more quickly and efficiently, comprehensively analyzes users' behavioral characteristics based on their network usage history, comprehensively analyzes users' behavioral characteristics, interest distributions, and the distributions of their personal existing knowledge fields through personalized computing for Internet users, obtains user characteristics through the analysis of the personal existing knowledge system of network users, and optimizes and guides the reflection of users' true search intentions during the use of search engines, the selection of professional products in online shopping, and the quality of the generated text information.
[0015] To achieve the above technical effects, the technical solutions adopted by the present invention are as follows:
[0016] The personalized knowledge and interest profiling system for Internet users obtains user characteristics through the analysis of the personal existing knowledge system of network users. Based on the analysis of Internet user data, an Internet user knowledge and interest profiling system based on an improved LDA model is proposed. On the basis of the improved LDA model analyzing users' interests, the knowledge distribution of users is further profiled;
[0017] The present invention expands from the level of user interest description, proposes a method for user description by profiling according to both knowledge and interest aspects to precisely profile users; incorporates the concept of the knowledge level of Internet users into the improved LDA model, and proposes a personalized knowledge and interest profiling system for Internet users; in user interest modeling, a keyword model is mainly used to model users' interests, and in user knowledge modeling, the knowledge level is incorporated into the user knowledge profiling model, and a directed connected graph is used to model users' knowledge;
[0018] There are three levels of user analysis for the personalized profiling model of Internet users' knowledge and interests as follows:
[0019] At the first level, for each user v, there are interests and existing knowledge corresponding to multiple topics, generating a topic distribution x, a user-interest distribution j, and a user-knowledge distribution w;
[0020] At the second level, for each microblog q, there are two aspects: an interest measure and an existing knowledge measure corresponding to a topic, generating an interest-microblog distribution e j , a knowledge-microblog distribution e w ;
[0021] At the third level, for each word k, there are multiple topics corresponding to it, generating a topic-word distribution s;
[0022] The generation process of the personalized profiling probability model of Internet users' knowledge and interests is as follows:
[0023] Process 1, sampling distributions: G v,x -Dirichlet(B), L v,w -Dirichlet(T), F v,j -Dirichlet(E), R v,w -Dirichlet(S);
[0024] Process 2, for each user v = 1, 2,..., V,
[0025] Sampling distributions: H v -Dirichlet(A v ), H-Binomial(π);
[0026] Sampling distribution: j v -Dirichlet(F v );
[0027] Sampling distribution: w v -Dirichlet(R v );
[0028] Process 3, for each microblog q = 1, 2,..., Q,
[0029] Sampling topic: x-Multinomial(H v );
[0030] Sampling the depth type z of the microblog description topic - Multinomial(w v );
[0031] Process 4, for each word k = 1, 2,..., K,
[0032] Sampling word: k n,m -Multinomial(G v,x );
[0033] Description of user knowledge and interest profiling model parameters: A v , B, T, E, S are hyperparameters and Dirichlet prior distributions; π is a hyperparameter and Binomial prior distribution; H is a |V| * |X| dimensional matrix representing the user-topic distribution; G is a |V| * |X| * |K| dimensional matrix representing the user-topic-word distribution; R is a |V| * |X| * |W| dimensional matrix representing the user-topic-knowledge distribution; F is a |V| * |X| * |J| dimensional matrix representing the user-topic-interest distribution; L is a |V| * |W| * |Z| dimensional matrix representing the user-knowledge-hierarchy distribution; v, j, x, w, z, k are variable instances: user v, user interest j, topic x, existing knowledge w, knowledge level z, word k; V, J, X, W, Z, K represent the user set, interest set, topic set, knowledge set, hierarchy set, and word set respectively.
[0034] For the personalized profiling system of Internet users' knowledge and interests, further, constructing a personalized model of user knowledge and interests includes constructing an Internet user interest model and an Internet user knowledge model;
[0035] The Internet user interest modeling adopts a keyword-based user interest expression model, presenting a user's interests with a keyword list, obtaining keywords to represent a user in this way, and obtaining a user's keyword set through various mining techniques or machine learning methods.
[0036] For the personalized profiling system of Internet users' knowledge and interests, further, in constructing the Internet user knowledge model, from the perspective of the hierarchy of subject demand knowledge, knowledge is divided into four levels: survival knowledge, skill knowledge, spiritual knowledge, and self-actualization knowledge. The knowledge organization meets the need of subjectivizing human objective knowledge and implements a series of ordering organization activities for the disordered state of objective knowledge;
[0037] The prerequisite for knowledge organization is to maximize the discovery of objective knowledge. An important part of knowledge discovery is the discovery of knowledge structure. Knowledge structure is the induction of objective knowledge through specific methods to make it relevant. Knowledge structure is a network structure, which consists of two elements: the nodes of numerous knowledge factors and the connection of knowledge associations. Knowledge factors are meta-knowledge, and the structure expressing the association relationship between meta-knowledge becomes a knowledge chain. Meta-knowledge is the most basic unit that composes the knowledge network. Each independently existing meta-knowledge is associated through a knowledge chain. As the knowledge system expands, knowledge only increases and does not disappear. Continuously discovering new meta-knowledge and continuously expressing the associations between meta-knowledge are the starting point and beginning for carrying out other activities of knowledge management;
[0038] The process of the subject discovering meta-knowledge and its association relationships is a gradual one. According to the specific objective environment, knowledge is organized, and the discovered meta-knowledge is processed through methods such as knowledge expression, knowledge recombination, knowledge storage and retrieval, knowledge clustering, knowledge layout, knowledge editing, and knowledge monitoring, ultimately improving the degree of knowledge orderliness and enhancing knowledge centralization.
[0039] For the personalized profiling system of Internet users' knowledge and interests, further, the expression of users' knowledge levels: The knowledge levels of Internet users present a network structure, and at the same time, there are clear hierarchical relationships in this network structure. The present invention uses a directed connected graph to represent this network structure, sets a super-root root node to connect different fields, and the nodes of these fields diverge from the super-root. Under these different fields, there are subdivided sub-fields, only expanding to the next-level nodes under the sub-fields, and the words at the fourth level will also belong to different sub-fields at the same time, forming a four-level directed connected graph composed of super-root - field - sub-field - word to represent the knowledge hierarchical structure of users; In this hierarchical structure, the more the knowledge fields covered by a user tend to the bottom layer of this knowledge hierarchical graph, the closer this user is to the expert level in this field. In this way, the knowledge depth of a user is represented.
[0040] For the personalized profiling system of Internet users' knowledge and interests, further, the analysis of Internet platform users is to conduct in-depth data mining on the text traces generated by Internet users' online activities and use of mobile APPs to obtain an objective and personalized understanding of users. The present invention uses the text data generated by the mobile APP of Sina Weibo as the data set for analysis;
[0041] User descriptions are classified into two categories: objective user descriptions and subjective user descriptions. Subjective user descriptions include the internal expressions of user interests, existing knowledge distributions, and behavioral patterns. Among these, two fields, namely user interests and user knowledge, are deeply explored. Users simultaneously apply user interest vectors and existing knowledge vectors, and the corresponding attribute dimensions of these two vectors are the same. Each attribute dimension represents the same event class or topic. The values corresponding to each dimension in the user interest vector and the existing knowledge vector represent the degree of interest and the depth of understanding of the user for the event class or topic represented by that dimension.
[0042] The text generation is represented in the form of a bag of words. There is a certain probability distribution for the theme or topic. The method of reproducing the text generation is through a probability model, and then the theme is obtained by combining parameter estimation. When the specified number of themes is a value M, after continuous supervised or unsupervised training and learning of the theme model, M themes are obtained, and the entire model is constructed based on the topic model.
[0043] The personalized profiling system for the knowledge and interests of Internet users. Further, the Internet user interest profiling architecture has the following three levels:
[0044] Level 1: For each user v, there are multiple themes and interests, generating a theme distribution x and a user-interest distribution j;
[0045] Level 2: For each microblog q, there is an aspect of theme interest measurement, generating an interest-microblog distribution e j ;
[0046] Level 3: For each word k, the theme it corresponds to generates a theme-word distribution s;
[0047] The generation process of the interest profiling probability graph model for Internet users is as follows:
[0048] First, the sampling distribution: G v,x -Dirichlet(B), F v,j -Dirichlet(E);
[0049] Second, for each user v = 1, 2,..., V,
[0050] The sampling distribution: H v -Dirichlet(A v ),H-Binomial(π);
[0051] The sampling distribution: j v -Dirichlet(F v );
[0052] Third, for each microblog q = 1, 2,..., Q,
[0053] Sampling topic: x-Multinomial(H v );
[0054] Fourth, for each word k = 1, 2,..., K,
[0055] Sampling word: k n,m -Multinomial(G v,x );
[0056] Description of user interest profiling model parameters: A v , B, and E are hyperparameters and Dirichlet prior distributions; π is a hyperparameter and Binomial prior distribution; H is a |V| * |X| dimensional matrix representing the user-topic distribution; G is a |V| * |X| * |K| dimensional matrix representing the user-topic-word distribution; F is a |V| * |X| * |J| dimensional matrix representing the user-topic-interest distribution; v, j, x, k are variable instances: user v, user interest j, topic x, word k; V, J, X, K represent the user set, interest set, topic set, and word set respectively.
[0057] The personalized profiling system for the knowledge and interests of Internet users. Further, the user analysis of the Internet user knowledge profiling architecture has the following three levels:
[0058] The first level, for each user v, corresponding to multiple topics and related existing knowledge, generates a topic distribution x and a user-knowledge distribution w;
[0059] The second level, for each microblog q, corresponding to a measure of the existing knowledge degree of a topic, generates a knowledge-microblog distribution e w ;
[0060] The third level, for each word k, corresponding to multiple topics and event classes, generates a topic-word distribution s;
[0061] The generation process of the Internet user knowledge profiling analysis probability model is as follows:
[0062] Step 1, sampling distribution: G v,x -Dirichlet(B), L v,w -Dirichlet(T), R v,w -Dirichlet(S);
[0063] Step 2, for each user v = 1, 2,..., V,
[0064] Sampling distribution: H v -Dirichlet(A v ), H - Binomial(π);
[0065] Sampling distribution: j v -Dirichlet(R v );
[0066] Step 3. For each Weibo q = 1, 2,..., Q,
[0067] Sampling topic: x - Multinomial(H v );
[0068] Sampling the depth type z of the topic described by the Weibo - Multinomial(w v );
[0069] Step 4. For each word k = 1, 2,..., K,
[0070] Sampling word: k n,m -Multinomial(G v,x );
[0071] Explanation of the parameters of the user knowledge profiling model: A v , B, T, S are hyperparameters and Dirichlet prior distributions; π is a hyperparameter and Binomial prior distribution; H is a |V| * |X| - dimensional matrix representing the user - topic distribution; G is a |V| * |X| * |K| - dimensional matrix representing the user - topic - word distribution; R is a |V| * |X| * |J| - dimensional matrix representing the user - topic - knowledge distribution; L is a |V| * |W| * |Z| - dimensional matrix representing the user - knowledge - level distribution; v, x, w, z, k are variable instances: user v, topic x, existing knowledge w, knowledge level z, word k; V, X, W, Z, K represent the user set, topic set, knowledge set, level set, and word set respectively.
[0072] The knowledge and interest personalized profiling system for Internet users. Further, the user personalized profiling model classifies common events and summarizes the daily life topics that the user is interested in. The fields described by the user profiling are divided into event classification and topic set, expressed as a feature set: {{Event class set} n1 ,{Topic set} n2} where n1 + n2 = n. For a specific user, both the user's interest and existing knowledge correspond to n - dimensional vectors with the feature set as the feature dimension. The feature set reflects the generality of the Internet user group, and the specific feature vector reflects the particularity of the Internet user individual;
[0073] Analysis of the user's existing knowledge, combined with the user's self-labeled description, prior knowledge, and in-depth analysis of the user's views and commentary on events or topics. The specific views and commentary depth are determined based on the content depth of the topics and events in the user's published text. Hierarchical classification is performed on the event description words and terms. A probabilistic graphical model is used to analyze the user's interests and existing knowledge. In the topic model of the present invention, a binomial distribution variable π is added to distinguish between the two. The level of the domain words used in the user's text in the knowledge domain tree is used to distinguish the depth of event analysis. The depth of the domain words for an event or topic is divided into five levels. A multinomial distribution variable z is introduced to identify the topic-word depth level distribution and the event-class-word depth level distribution.
[0074] The personalized profiling system for the knowledge and interests of Internet users. Further, the analysis data input into the model is the user's Weibo data. At the same time, each user corresponds to the user's own Weibo text set. The Weibo texts of these M users are used as the input data set of the model. The results have two parts, one is the user's interests, and the other is the user's knowledge.
[0075] Regarding the user's interests, the model obtains the vector matrix of each user's Weibo text and interests, and calculates the weight of each user for each interest field. The calculation formula is:
[0076]
[0077] In the formula, x v,j represents the weight of user v for the jth interest. N represents the number of Weibo texts of user v in the text set. k v,w,j represents the weight of the jth interest in the wth Weibo text of user v. For N users, their corresponding interest weights are calculated. The interest vector of each user can be expressed as {x v,1 , x v,2 , x v,3 , …, x v,j , …};
[0078] Regarding the user's knowledge, all the user's Weibo texts are divided into four categories by the model: super root, super topic, sub topic, and word. The super root represents the root of all super topics and has no actual value in the model. The super topics include the interest-level fields of the user. The sub topics are the links between the super topics and the words and are the results of the subdivision of the super topics. The bottom-level words are segmented from the user's Weibo.
[0079] Compared with the prior art, the contributions and innovations of the present invention are:
[0080] First, the personalized profiling system for the knowledge and interests of Internet users provided by the present invention solves the problem that a vast amount of information brings great difficulties to rapid information search, meets the increasingly high requirements of users for information acquisition, provides precise personalized information search services, enables users to find accurate information that meets their needs more quickly and efficiently, comprehensively analyzes users' behavioral characteristics through their Internet usage history, comprehensively analyzes users' behavioral characteristics, interest distribution, and the distribution of their personal existing knowledge fields through personalized computing for Internet users, obtains user characteristics through the analysis of the personal existing knowledge system of network users, and optimizes and guides the reflection of users' true search intentions during the process of using search engines, the selection of professional products in online shopping, and the quality of generated text information;
[0081] Second, based on the analysis of the Weibo data of Internet users, the present invention proposes a profiling system for the knowledge and interests of Internet users based on an improved LDA model. On the basis of the improved LDA model analyzing users' interests, it further profiles the knowledge distribution of users, expands from the level of user interest description, and proposes a method for user description according to the profiling methods in terms of both knowledge and interests to accurately profile users personally; the concept of the knowledge level of Internet users is introduced and integrated into the improved LDA model. The user interest modeling mainly uses the keyword model to model users' interests, and the user knowledge modeling incorporates the knowledge level into the user's knowledge profiling model and models the user knowledge with a directed connected graph. The results show that the present invention can accurately and quickly profile users, has high feasibility and accuracy, and has been applied in actual Internet user personalization projects;
[0082] Third, the status of personalized services in Internet user profiling has begun to become prominent and is becoming increasingly important in the field of user interest mining. The key to the personalized services of the present invention lies in mining users' interests and hobbies. If it can accurately mine a user's interests and hobbies, the search engine will not have a large number of spam messages unrelated to the user's search target information under the same keyword. Instead, it will give the search results that best meet the user's needs based on the keyword and different users' interests and hobbies, maximizing the satisfaction of users' needs;
[0083] Fourth, the knowledge and interest personalized profiling system for Internet users provided by the present invention comprehensively analyzes the user behavior characteristics, interest distribution, and the distribution of the user's existing knowledge fields based on the Internet usage history of the Internet user, and creatively solves the problems of behavior representation, interest quantification, knowledge field measurement, and the combination points among the three. It can overall profile a real and specific user, and the returned search results will also feedback more in line with the user's intention based on the understanding of the user. Users no longer need to spend a large amount of unnecessary time to filter the still messy search results twice or even three times to find the information they really need, which not only saves the user's time but also saves the user's energy. For the user experience, it also plays a huge role in improvement and promotion, and can provide an intelligent search function that can provide a precise event summary of event elements that more meets the user's personalized search requirements and understands the user's logic;
[0084] Fifth, deeply analyzing the characteristics of Internet users, such as the user's interests and knowledge level, is the premise for correctly recommending potential purchase products to users. Users will not be submerged in a dazzling array of products that do not even meet the user's needs at all. In aspects such as the selection of professional products and the quality of generated text information in online shopping, it is a very crucial factor. The present invention is conducive to converting potential users in front of the mobile phone screen into actual consumers, recommending more products that they are interested in to website registered users and stimulating consumption. It is a key technology for gaining a foothold in the highly competitive e-commerce industry and has broad market application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 is a schematic diagram of the Internet user interest profiling architecture of the present invention.
[0086] Figure 2 is a schematic diagram of the Internet user interest profiling architecture of the present invention.
[0087] Figure 3 is a schematic diagram of the personalized profiling architecture of the knowledge and interests of Internet users. DETAILED DESCRIPTION OF THE INVENTION
[0088] The following combines the drawings to further describe the technical solutions of the knowledge and interest personalized profiling system for Internet users provided by the present invention, so that those skilled in the art can better understand the present invention and be able to implement it.
[0089] With the rapid development of the mobile Internet, the amount of data has expanded rapidly, and the resulting massive information has brought great difficulties to rapid information search. At the same time, users' requirements for information acquisition are also getting higher and higher. More and more users are not satisfied with simple information search services and need to quickly and efficiently find accurate information that meets their needs. This has brought new challenges and opportunities to personalized data services. An important part of personalized data services is to comprehensively analyze users' behavioral characteristics through users' network usage traces and comprehensively analyze users' behavioral characteristics, interest distributions, and the distribution of users' existing knowledge fields through personalized computing of Internet users.
[0090] The prior art does not pay enough attention to the existing knowledge distribution of users. The present invention obtains user characteristics through the analysis of the personal existing knowledge system of network users, optimizes and guides the reflection of the true search intention during the process of users using search engines, the selection of professional products in online shopping, and the quality of generated text information. Based on the analysis of the microblog data of Internet users, a system for depicting the knowledge and interests of Internet users based on an improved LDA model is proposed. On the basis of the improved LDA model analyzing users' interests, the knowledge distribution of users is further depicted.
[0091] The present invention expands from the level of user interest description, proposes a method for user description according to the description methods of knowledge and interest, and accurately depicts users' personalities; introduces the concept of the knowledge level of Internet users, integrates the knowledge level into the improved LDA model, and proposes a system for personalized description of the knowledge and interests of Internet users; in user interest modeling, a keyword model is mainly used to model users' interests, and in user knowledge modeling, the knowledge level is incorporated into the user's knowledge description model, and a directed connected graph is used to model users' knowledge.
[0092] Through examples, the present invention analyzes the microblogs of different people in ten fields, and depicts these people in these two aspects on the basis of the system for personalized description of the knowledge and interests of Internet users. The results show that the present invention can accurately and quickly depict users' characteristics, has high feasibility and accuracy, and has been applied in actual Internet user personalization projects.
[0093] I. Construct a personalized model of user knowledge and interest
[0094] The status of personalized services in depicting Internet users has begun to stand out and become increasingly important in the field of user interest mining. Under the bombardment of complex Internet information, everyone wants to obtain the most important knowledge content for themselves at the fastest speed. The fastest way is to use a search engine. However, if the search engine returns the same results for every keyword searched by each person, it will only make users spend more time on information filtering. Such a search engine does not provide the information that users most want according to their interests and hobbies. The key to personalized services lies in mining users' interests and hobbies. If the interests and hobbies of a user can be accurately mined, then the search engine will not have a large amount of junk information unrelated to the user's search target information under the same keyword. Instead, it will give the search results that best meet the user's needs based on the keyword and the interests and hobbies of different users, maximizing the satisfaction of user needs.
[0095] (1) Construct an Internet user interest model
[0096] 1. Internet user interest modeling
[0097] With the continuous development and maturity of Internet personalized services, to gradually improve and fully realize personalized services, a model is needed to comprehensively depict a series of relevant information such as users' interests, hobbies, and habits.
[0098] The user modeling of the present invention does not require the user himself to participate in the modeling process. During the system modeling process, the user does not need to subjectively provide any information related to the user himself. The information required by the system is implicitly obtained and parsed by the system. The user modeling method does not require the participation of the user, and the obtained information has more real and reliable data compared with interactive user modeling, and the user experience is also better.
[0099] 2. Expression of the user interest model
[0100] The user modeling correctly reflects the information about users' interests and hobbies, which is based on the expression of the user interest model. The present invention adopts a keyword user interest expression model, presenting a user's interests with a keyword list. For example, if a user is interested in electronic products, it can be expressed with keywords as follows {tablet, mobile phone, smart bracelet, smart watch, electronic device}. Many keywords are provided when the user registers, and these keywords are selected by themselves. In this way, obtaining keywords can represent a user to a certain extent. This is an example method. The present invention obtains a user's keyword set through various mining technologies or machine learning methods.
[0101] (2) Construct an Internet user knowledge model
[0102] 1. Internet user knowledge hierarchical structure
[0103] From the hierarchical perspective of the knowledge of the subject's needs, knowledge is divided into four levels: survival knowledge, skill knowledge, spiritual knowledge, and self-actualization knowledge. Knowledge organization meets the need of subjectivizing human objective knowledge and implements a series of ordering organization activities for the disordered state of objective knowledge.
[0104] The prerequisite for knowledge organization is to maximize the discovery of objective knowledge. Knowledge refers to the information processed by the human brain and exists attached to different types of carriers. The existence of knowledge is not isolated but shows systematic and structural characteristics. An important part of knowledge discovery is the discovery of the knowledge structure. The knowledge structure is a scientific and reasonable induction of objective knowledge through specific methods to make it relevant. The structure of knowledge is a network structure, which consists of two elements: the nodes of numerous knowledge factors and the node connections of knowledge associations. Knowledge factors are meta-knowledge, and the structure expressing the association relationship between meta-knowledge becomes the knowledge chain. Meta-knowledge is the most basic unit that makes up the knowledge network. Each independently existing meta-knowledge is associated through the knowledge chain. The knowledge system objectively exists and does not change due to the understanding of the subject, which is the objective existence of the knowledge system. In addition, the knowledge system also has the characteristics of gradual increase and indestructibility, that is, its system expands continuously according to the changes of the times and the environment. Knowledge only increases and does not disappear. Continuously discovering new meta-knowledge and continuously expressing the associations between meta-knowledge are the starting point and beginning of carrying out other activities of knowledge management.
[0105] The process of the subject discovering meta-knowledge and its association relationship is a gradual one. According to the specific objective environment, knowledge organization is carried out. The discovered meta-knowledge is processed through methods such as knowledge expression, knowledge recombination, knowledge storage and retrieval, knowledge clustering, knowledge layout, knowledge editing, and knowledge monitoring, and finally the degree of orderliness of knowledge is improved and the concentration of knowledge is enhanced.
[0106] 2. User Knowledge Hierarchical Expression
[0107] The knowledge level of Internet users presents a network structure, and there are also clear hierarchical relationships in this network structure. The present invention uses the method of a directed connected graph to represent this network structure. Since the upper knowledge level has a large coverage range and is divided into many different fields according to different aspects, some of these fields have intersections and some do not. Therefore, a super root root node is set to connect different fields. The nodes of these fields diverge from the super root. Under these different fields, there will be subdivided sub-fields. The sub-fields under different fields may be the same. For the convenience of parsing data, only expand to the nodes of the next layer of the sub-fields. The words in the fourth layer will also belong to different sub-fields at the same time. In this way, a four-layer directed connected graph composed of super root - field - sub-field - word is formed to represent the knowledge hierarchical structure of users.
[0108] In this hierarchical structure, the more the knowledge fields covered by a user tend to the bottom layer of this knowledge hierarchy diagram, the closer this user is to the expert level in this field. In this way, the knowledge depth of a user is represented.
[0109] II. User Profiling Model
[0110] The analysis of Internet platform users conducts in-depth data mining on the text traces generated by Internet users' online activities and their use of mobile APPs to obtain an objective and personalized understanding of users. Based on the data obtained through the analysis, data and application push, knowledge dissemination, and various different purposes of network marketing are carried out. Since the data volume generated by the Weibo APP is very rich and close to the most authentic and reliable information expressed by Internet users, the present invention uses the text data generated by the mobile APP of Sina Weibo as the data set for analysis.
[0111] User descriptions are classified into two categories: user objective descriptions and user subjective descriptions. User objective descriptions include objective and factual descriptions of the user's gender, age, height, native place, education level, and marital status. User subjective descriptions include the internal expressions of the user's interests, existing knowledge distribution, and behavior patterns. Among them, two fields, namely user interests and user knowledge, are deeply mined. The user corresponds to a user interest vector and an existing knowledge vector at the same time, and the attribute dimensions corresponding to these two vectors are the same. Each attribute dimension represents the same event class or topic. The values corresponding to each dimension in the user interest vector and the existing knowledge vector represent the user's interest degree and understanding depth of the event class or topic represented by this dimension.
[0112] The text generation is represented in the form of a bag of words, and there is a certain probability distribution for the theme or topic. On this basis, the text generation method is reproduced through a probability model, and then the theme is obtained by combining parameter estimation. When the specified number of themes is a value M, through continuous supervised or unsupervised training and learning of the theme model, M themes are obtained. Combining the characteristics of the Internet user platform, the present invention constructs the entire model based on the topic model.
[0113] (I) Internet User Interest Profiling Architecture
[0114] Interest is the attitude and emotion of an individual towards specific things, activities, and objects considered. It is an intangible driving force, and interests will be affected by different factors. The interests and hobbies of different people with different surrounding factors are different. The user interest profiling architecture model of the present invention is as Figure 1 shown.
[0115] The Internet user interest profiling architecture has the following three levels:
[0116] Level 1: For each user v, there are multiple corresponding themes and interests, generating a theme distribution x and a user-interest score j;
[0117] At the second level, for each microblog q, corresponding to a topic interest measurement aspect, the interest-microblog distribution e will be generated j ;
[0118] Level 3: For each word k, its corresponding topic will generate a topic-word distribution s;
[0119] The generation process of the probability graph model of Internet users’ interest description is as follows:
[0120] First, the sampling distribution: G v,x -Dirichlet(B), F v,j -Dirichlet(E);
[0121] Second, for each user v=1,2,...,V,
[0122] Sampling distribution: H v -Dirichlet(A v ), H-Binomial(π);
[0123] Sampling distribution: j v -Dirichlet(F v );
[0124] Third, for each microblog q=1,2,...,Q,
[0125] Sampling theme: x-Multinomial(H v );
[0126] Fourth, for each word k=1,2,...,K,
[0127] Sampling word: k n,m -Multinomial(G v,x );
[0128] User interest description model parameter description: A v , B, E are hyperparameters and Dirichlet prior distribution; π is a hyperparameter and Binomial prior distribution; H is a |V|*|X|-dimensional matrix, representing the user-topic distribution; G is a |V|*|X|*|K|-dimensional matrix, representing the user-topic-word distribution; F is a |V|*|X|*|J|-dimensional matrix, representing the user-topic-interest distribution; v, j, x, k are variable instances: user v, user interest j, topic x, word k; V, J, X, K represent the user set, interest set, topic set, and word set respectively.
[0129] (2) Internet User Knowledge Description Framework
[0130] Knowledge is the objective knowledge scope that an individual understands about specific things, activities, and objects. When a user faces a certain thing or activity, they combine their existing understanding of them in their minds to form a series of relevant information reserves about these things and activities. Knowledge cannot be the same for people in different environments, with different experiences, and different occupations. The user knowledge profiling architecture model is as shown in Figure 2 as follows.
[0131] There are three levels of user analysis for the Internet user knowledge profiling architecture as follows:
[0132] The first level: For each user v, there are multiple corresponding topics and relevant existing knowledge, generating a topic distribution x and a user-knowledge distribution w;
[0133] The second level: For each microblog q, there is a degree measure of existing knowledge for a corresponding topic, generating a knowledge-microblog distribution e w ;
[0134] The third level: For each word k, there are multiple corresponding topics and event categories, generating a topic-word distribution s;
[0135] The generation process of the Internet user knowledge profiling analysis probability model is as follows:
[0136] Step 1, sampling distributions: G v,x -Dirichlet(B), L v,w -Dirichlet(T), R v,w -Dirichlet(S);
[0137] Step 2, for each user v = 1, 2,..., V,
[0138] Sampling distributions: H v -Dirichlet(A v ), H - Binomial(π);
[0139] Sampling distribution: j v -Dirichlet(R v );
[0140] Step 3, for each microblog q = 1, 2,..., Q,
[0141] Sampling topic: x - Multinomial(H v );
[0142] Sampling the depth type z of the microblog description topic - Multinomial(w v );
[0143] Step 4, for each word k = 1, 2,..., K,
[0144] Sampling word: k n,m -Multinomial(G v,x )
[0145] Description of user knowledge profiling model parameters: A v , B, T, S are hyperparameters and Dirichlet prior distributions; π is a hyperparameter and Binomial prior distribution; H is a |V| * |X| dimensional matrix representing the user-topic distribution; G is a |V| * |X| * |K| dimensional matrix representing the user-topic-word distribution; R is a |V| * |X| * |J| dimensional matrix representing the user-topic-knowledge distribution; L is a |V| * |W| * |Z| dimensional matrix representing the user-knowledge-hierarchy distribution; v, x, w, z, k are variable instances: user v, topic x, existing knowledge w, knowledge level z, word k; V, X, W, Z, K represent the user set, topic set, knowledge set, hierarchy set, and word set respectively.
[0146] (III) Knowledge and Interest Personalized Profiling Model for Internet Users
[0147] User profiling mainly includes user interest profiling and user knowledge profiling. User interest refers to the preference degree of a user for a certain event category or topic category, and user knowledge refers to the understanding degree of the knowledge field of event categories and topics. The two belong to different categories. The existing technology does not pay due attention to user knowledge, and the important innovation point of the present invention lies in the profiling of user knowledge.
[0148] The user personalized profiling model classifies common events and summarizes the daily life topics that the user is interested in. This is different from the limitation in user interest that only corresponds to topics, and also different from the ambiguity of dividing events into public topics and daily topics. The present invention divides the field of user profiling into event classification and topic set, expressed as a feature set: {{event class set} n1 ,{topic set} n2} where n1 + n2 = n. For a specific user, both user interest and existing knowledge correspond to n-dimensional vectors with the feature set as the feature dimension. Due to the different internal and external factors of user personality, the combination of these two feature vectors varies from user to user. The feature set reflects the generality of the Internet user group, and the specific feature vector reflects the particularity of the Internet user individual.
[0149] What makes the present invention different from the existing technology in terms of user interest or user behavior pattern research is the analysis of the user's existing knowledge. The analysis of the user's existing knowledge needs to be further judged by combining the user's self-label description with prior knowledge and the user's in-depth analysis of the event or topic view and commentary content description. The specific view and commentary depth needs to be judged based on the content depth of the topic and event in the user's published text, so it is necessary to classify the event description words and terms in a hierarchical manner. The present invention uses a probabilistic graph model to analyze user interests and user's existing knowledge. In the topic model of the prior art, event classes and topics cannot be accurately distinguished. Therefore, a binomial distribution variable π is added to the topic model of the present invention to distinguish between the two; secondly, the level of the domain words used in the user text in the knowledge domain tree is distinguished for the depth of event analysis to facilitate the study of the user's existing knowledge. The depth of the domain words for events or topics is divided into five layers, and a multinomial distribution variable z is introduced to identify the topic-word depth hierarchical distribution and the event class-word depth hierarchical distribution. The personalized description architecture of user knowledge and interests is as follows: Figure 3 shown.
[0150] The user analysis of the personalized description model of Internet user knowledge and interests has the following three levels:
[0151] At the first level, for each user v, corresponding to the interests and existing knowledge of multiple topics, a topic distribution x, user-interest distribution j and user-knowledge distribution w are generated;
[0152] At the second level, for each microblog q, the interest metric and the existing knowledge level of a topic are used to generate the interest-microblog distribution e j , Knowledge-Weibo Distribution w ;
[0153] At the third level, for each word k, corresponding to multiple topics, a topic-word distribution s is generated.
[0154] The generation process of the probability model for personalized description and analysis of Internet users' knowledge and interests is as follows:
[0155] Process 1, sampling distribution: G v,x -Dirichlet(B), L v,w -Dirichlet(T), F v,j -Dirichlet(E), R v,w -Dirichlet(S);
[0156] Process 2: For each user v=1,2,...,V,
[0157] Sampling distribution: H v -Dirichlet(A v), H-Binomial(π);
[0158] Sampling distribution: j v -Dirichlet(F v );
[0159] Sampling distribution: w v -Dirichlet(R v );
[0160] Process three, for each microblog q = 1, 2,..., Q,
[0161] Sampling topic: x-Multinomial(H v );
[0162] Sampling the depth type z-Multinomial(w of the microblog describing the topic v );
[0163] Process four, for each word k = 1, 2,..., K,
[0164] Sampling word: k n,m -Multinomial(G v,x );
[0165] Description of the parameters of the user knowledge and interest profiling model: A v , B, T, E, S are hyperparameters and Dirichlet prior distributions; π is a hyperparameter and Binomial prior distribution; H is a |V|*|X| dimensional matrix representing the user-topic distribution; G is a |V|*|X|*|K| dimensional matrix representing the user-topic-word distribution; R is a |V|*|X|*|W| dimensional matrix representing the user-topic-knowledge distribution; F is a |V|*|X|*|J| dimensional matrix representing the user-topic-interest distribution; L is a |V|*|W|*|Z| dimensional matrix representing the user-knowledge-hierarchy distribution; v, j, x, w, z, k are variable instances: user v, user interest j, topic x, existing knowledge w, knowledge level z, word k; V, J, X, W, Z, K represent the user set, interest set, topic set, knowledge set, hierarchy set, and word set respectively.
[0166] III. Embodiment Design and Analysis
[0167] (I) Embodiment Dataset
[0168] In this embodiment, ten fields are selected, namely politics, economy, automotive, history, military, Internet, environment, real estate, art, and sports. Five high-quality users are selected from each of these ten fields. High-quality users refer to those who have a good understanding of the field and whose Weibo content is of relatively high quality. A total of 50 users are selected, and the Weibo data of these 50 users is crawled on Weibo. For each user, 1000 Weibo posts are crawled. The positioning of expert-level users means that these users are more deeply involved in this field.
[0169] (2) Data preprocessing
[0170] When crawling Weibo, certain conditions are added to restrict the Weibo. There are two specific forms in the user's Weibo. One is forwarding someone else's Weibo, and the other is the Weibo posted by the author himself. When processing the crawled Weibo, if it is a forwarded and commented Weibo, the words "Forwarded" and "@" in the middle are removed, and the user's comment is merged and associated with the Weibo he commented on. Forwarding someone else's Weibo reflects the user's interests and knowledge level.
[0171] In terms of word segmentation, the Chinese Lexical Analysis System ICTCLAS is adopted. This system has a high word segmentation accuracy rate, supports multiple encodings such as GBK, UTF-8, and BIG5, provides APIs for multiple platforms, and new Weibo word segmentation.
[0172] (3) Embodiment design
[0173] The analysis data input into the model is the user's Weibo data. At the same time, each user corresponds to the user's own Weibo text set. The Weibo texts of these M users are used as the input data set of the model. The results have two parts. One part is the user's interest, and the other part is the user's knowledge. In terms of user interest, the model obtains the vector matrix of each user's Weibo text and interest, and calculates the weight of each user for each interest field. The calculation formula is:
[0174]
[0175] In the formula, x v,j represents the weight of user v for the jth interest, N represents the number of Weibo texts in the Weibo text set of user v, k v,w,j represents the weight of the jth interest in the wth Weibo text of user v. For N users, their corresponding interest weights are calculated. The interest vector of each user can be expressed as {x v,1 , x v,2 , x v,3 , …, x v,j , …};
[0176] In terms of user knowledge, all the user's microblogs are classified into four categories by the model: super root, super topic, subtopic, and word. The super root represents the root of all super topics and has no actual value in the model. Super topics include the user's interest areas. Subtopics are the link between super topics and words and are the results of the subdivision of super topics. The bottom-level words are segmented from the user's microblogs. In the model, the number of super topics is set to 1, the number of subtopics is set to 10, the number of words under each subtopic is 5, and there are 50 words at the bottom layer of this hierarchy. The user knowledge hierarchy diagram is generated based on all the data sets in the embodiment and is not specific to a particular user. On the basis of obtaining the user's interest distribution, the words in the user's microblogs are combined with the knowledge hierarchy distribution to obtain the characterization of the user's knowledge level. The knowledge depth of user v in a specific interest area is specifically calculated by the following formula:
[0177]
[0178]
[0179] dep(w) v represents the knowledge depth of user v, Z represents the depth of the knowledge hierarchy. According to the above model, Z = 4; weight j represents the weight of the node in the j-th layer. In this embodiment, the weight is assigned as: weight j = e j-1 ; N represents the total number of the user's microblogs; f(j) represents the number of words or topics in the j-th layer that appear in the microblog text set of user v. For repeated words or topics, they are only counted once; show j,i represents whether the word or topic in the j-th layer of the hierarchical structure appears in the microblog text set of user v. If it appears, the value is 1; if not, the value is 0.
[0180] (IV) Analysis of the Embodiment
[0181] Based on the above steps, set the threshold of the user interest weight to 0.1. As long as the user's interest weight value exceeds 0.1 in certain fields, classify it into the user's interests. For the definition of user knowledge, in the model, since there are too many nodes in the hierarchy, randomly select several nodes to represent a simple hierarchical structure. It can be seen that there are two super-topics: Internet and Automobile. Under these two topics of Internet and Automobile, there is also a common sub-topic: Internet of Vehicles. Under the super-topic of Internet, there are sub-topics such as big data, and there are more professional terms under the distributed system. The same is true under the topic of automobile structure. In such a hierarchical structure, the more towards the bottom layer, the more professional the vocabulary involved. If a user's Weibo posts involve more professional vocabulary, to a certain extent, it can confirm the depth of their knowledge in this field, and the deeper the knowledge, the more accurately the user's knowledge in this field can be depicted. In this embodiment, users are divided into five levels according to the depth of knowledge, namely: General, Understanding, Comprehension, Mastery, and Professional. These five levels are divided according to the user's dep(w) value.
[0182] Knowledge is hidden in interests. The model mines the depth of user knowledge based on interests to confirm whether a user only stays at the superficial level of being interested in a certain interest or further belongs to the depth of mastery or even professional cognitive level. In terms of interest characterization, mainly compare the user's interests with the user's own labels. If the user's interests can be covered by the labels defined by the user themselves and the number of inconsistent cases is less than or equal to one label, it is determined to be accurate; otherwise, it is inaccurate. In terms of knowledge characterization evaluation, if the difference between the manually marked result and the model result is one level, that is, one is level five and the other is level four, it is still determined that the result is accurate. However, once the difference reaches or is greater than two levels, it is determined that the result is inaccurate.
[0183] The embodiment compares and statistically analyzes the results obtained by running 50 selected high-quality users in the user characterization model with the average discrimination results of manual marking. The results show that in terms of interest characterization, the comparison results of 43 users are accurate, with an accuracy rate as high as 86%; in terms of knowledge characterization, the comparison results of 42 users are accurate, with an accuracy rate of 84%. This indicates that the present invention is relatively accurate in user characterization. After careful research, it is found that since a large number of Weibo posts forwarded by some users contain data such as pictures and long Weibo posts that the model cannot read and process, and these pictures and long Weibo posts are taken into account during manual marking, to a certain extent, it reduces the accuracy rate of the model.
[0184] With the rapid development of the Internet, the amount of network information data is increasing, and people's requirements for quickly and accurately obtaining information are getting higher and higher. Personalized services are also facing increasing challenges, and it is necessary to accurately depict users' personalities to improve their own service quality. Based on Internet users and combined with users' Weibo information, the present invention has made the following innovations: First, the present invention is not limited to the level of interest description, and a method for depicting users is proposed based on the description methods of knowledge and interest. Second, the knowledge level is integrated into the improved LDA model, and a personalized description model of knowledge and interest of Internet users is proposed. The results of the embodiments show that the present invention can relatively accurately depict the knowledge and interest of Internet users.
Claims
1. A personalized profiling system for the knowledge and interests of Internet users, characterized in that, User characteristics are obtained by analyzing the personal existing knowledge system of network users. Based on the analysis of Internet user data, an Internet user knowledge and interest profiling system based on an improved LDA model is proposed. On the basis of the improved LDA model analyzing users' interests, the knowledge distribution of users is further profiled; From the level of user interest description, a method of user description is proposed according to the profiling methods of knowledge and interest, so as to accurately conduct personalized characterization of users; The concept of Internet user knowledge level is added, and the knowledge level is integrated into the improved LDA model, and an Internet user knowledge and interest personalized profiling system is proposed; In user interest modeling, a keyword model is used to model user interests. In user knowledge modeling, the knowledge level is incorporated into the user's knowledge profiling model, and a directed connected graph is used to model user knowledge; User interest model expression: It reflects information on users' interests and hobbies. The keyword user interest expression model is adopted, and a keyword list is used to present a user's interests. A keyword set of a user is obtained through mining technology or machine learning methods; Internet user knowledge hierarchy structure: From the perspective of the hierarchy of subject demand knowledge, knowledge is divided into four levels: survival knowledge, skill knowledge, spiritual knowledge, and self-actualization knowledge. A series of orderly organization activities are carried out for the disordered state of objective knowledge; The structure of knowledge is a network structure, which consists of two elements: the nodes of numerous knowledge factors and the node connections of knowledge associations. Knowledge factors are meta-knowledge, and the structure expressing the association relationship between meta-knowledge becomes a knowledge chain. Meta-knowledge is the most basic unit of the knowledge network. Each independently existing meta-knowledge is associated through a knowledge chain, continuously expressing the association between meta-knowledge; The subject gradually discovers meta-knowledge and its association relationships, organizes knowledge according to the specific objective environment, and finally improves the degree of knowledge order through methods such as knowledge expression, knowledge recombination, knowledge storage and retrieval, knowledge clustering, knowledge layout, knowledge editing, and knowledge monitoring; User knowledge level expression: The knowledge level of Internet users presents a network structure, and at the same time, there are clear hierarchical relationships in the network structure. A directed connected graph is used to represent this network structure. It is divided into many different domains according to different aspects, and a super root root node is set to connect different domains. The nodes of these domains diverge from the super root. Under these different domains, sub-domains are subdivided, and only extended to the next-level nodes of the sub-domains. The words at the fourth level will also belong to different sub-domains at the same time, constituting a four-level directed connected graph composed of super root - domain - sub-domain - word to represent the user's knowledge level structure; In this hierarchical structure, the more the knowledge domains covered by a user tend to the bottom of this knowledge level graph, the closer this user is to the expert level in this domain. In this way, the knowledge depth of a user is represented; User Tracing Model: Internet platform user analysis deeply mines the text traces generated by Internet users' online activities and their use of mobile APPs to obtain an objective and personalized understanding of users. Based on the data obtained from the analysis, targeted data and application push, knowledge dissemination, and online marketing are carried out. The text data generated by the mobile APP of Sina Weibo is used as the dataset for analysis; User descriptions are classified into two categories: objective user descriptions and subjective user descriptions. Objective user descriptions include the objective facts of the user's gender, age, height, native place, education level, and marital status. Subjective user descriptions include the internal expressions of the user's interests, existing knowledge distribution, and behavior patterns. Among them, two fields, namely user interests and user knowledge, are deeply mined. The user corresponds to a user interest vector and an existing knowledge vector at the same time, and the attribute dimensions corresponding to these two vectors are the same. Each attribute dimension represents the same event class or topic. The values corresponding to each dimension in the user interest vector and the existing knowledge vector represent the user's interest degree and understanding depth of the event class or topic represented by this dimension; The text generation is represented in the form of a bag of words. There is a certain probability distribution for the theme or topic. On this basis, the text generation method is reproduced through a probability model, and then the theme is obtained by combining parameter estimation. When the specified number of themes is a value M, through continuous supervised or unsupervised training and learning of the theme model, M themes are obtained. Combining the characteristics of the Internet user platform, the entire model is constructed based on the topic model; Internet User Interest Tracing Architecture: Level 1: For each user v, there are multiple corresponding themes and interests, generating a theme distribution x and a user-interest distribution j; Level 2: For each Weibo q, there is a theme interest measurement aspect, generating an interest-Weibo distribution ej; Level 3: For each word k, the theme it corresponds to generates a theme-word distribution s; The generation process of the interest tracing probability graph model for Internet users is as follows: First, sampling distribution: Gv,x - Dirichlet(B), Fv,j - Dirichlet(E); Second, for each user v = 1, 2,..., V, Sampling distribution: Hv - Dirichlet(Av), H - Binomial(π); Sampling distribution: jv - Dirichlet(Fv); Third, for each Weibo q = 1, 2,..., Q, Sampling theme: x - Multinomial(Hv); Fourth, for each word k = 1, 2,..., K, Sampling word: kn,m - Multinomial(Gv,x); User interest profiling model parameter description: Av, B, and E are hyperparameters and Dirichlet prior distributions; π is a hyperparameter and Binomial prior distribution; H is a |V| * |X| dimensional matrix representing the user-topic distribution; G is a |V| * |X| * |K| dimensional matrix representing the user-topic-word distribution; F is a |V| * |X| * |J| dimensional matrix representing the user-topic-interest distribution; v, j, x, k are variable instances: user v, user interest j, topic x, word k; V, J, X, K represent the set of users, the set of interests, the set of topics, and the set of words respectively; Internet user knowledge profiling framework: At the first level, for each user v, there are multiple topics and related existing knowledge, generating a topic distribution x and a user-knowledge distribution w; At the second level, for each microblog q, there is a measure of the degree of existing knowledge of a topic, generating a knowledge-microblog distribution ew; At the third level, for each word k, there are multiple topics and event classes, generating a topic-word distribution s; The generation process of the Internet user knowledge profiling analysis probability model is as follows: Step 1, sampling distributions: Gv,x - Dirichlet(B), Lv,w - Dirichlet(T), Rv,w - Dirichlet(S); Step 2, for each user v = 1, 2,..., V, Sampling distributions: Hv - Dirichlet(Av), H - Binomial(π); Sampling distribution: jv - Dirichlet(Rv); Step 3, for each microblog q = 1, 2,..., Q, Sampling topic: x - Multinomial(Hv); Sampling the depth type z of the microblog describing the topic - Multinomial(wv); Step 4, for each word k = 1, 2,..., K, Sampling word: kn,m - Multinomial(Gv,x) User knowledge profiling model parameter description: Av, B, T, S are hyperparameters and Dirichlet prior distributions; π is a hyperparameter and Binomial prior distribution; H is a |V| * |X| dimensional matrix representing the user-topic distribution; G is a |V| * |X| * |K| dimensional matrix representing the user-topic-word distribution; R is a |V| * |X| * |J| dimensional matrix representing the user-topic-knowledge distribution; L is a |V| * |W| * |Z| dimensional matrix representing the user-knowledge-level distribution; v, x, w, z, k are variable instances: user v, topic x, existing knowledge w, knowledge level z, word k; V, X, W, Z, K represent the set of users, the set of topics, the set of knowledge, the set of levels, and the set of words respectively; Improvement of user knowledge depth hierarchical description: The user personalized description model classifies common events and summarizes daily life topics that users are interested in. The user description domain is divided into event categories and topic sets, which are represented as feature sets: {{event class set}} n1 ,{topic collection} n2 Where n1 + n2 = n. For a specific user, both user interests and existing knowledge correspond to n-dimensional vectors with feature sets as feature dimensions. The feature sets reflect the generality of the Internet user group, while the specific feature vectors reflect the particularity of individual Internet users. Based on the analysis of the user's existing knowledge, further judgment is made by combining the user's self-labeled description with prior knowledge and the in-depth analysis of the user's views and commentary on events or topics. The specific views and commentary depth are judged according to the content depth of the relevant topics and events in the user's published text, and the event description words and terms are hierarchically classified; a probabilistic graphical model is used to analyze the user's interests and existing knowledge, and a binomial distribution variable π is added to the topic model to distinguish the two; secondly, the level of the domain words used in the user's text in the knowledge domain tree is used to distinguish the depth of event analysis, which helps to study the user's existing knowledge. The depth of the domain words for an event or topic is divided into five levels, and a multinomial distribution variable z is introduced to identify the topic-word depth level distribution and the event-class-word depth level distribution; There are three levels of user analysis in the personalized profiling model of Internet users' knowledge and interests: The first level, for each user v, corresponding to the interests and existing knowledge of multiple topics, generating a topic distribution x, a user-interest distribution j, and a user-knowledge distribution w; At the second level, for each microblog q, an interest-microblog distribution e is generated in terms of the interest measure of a corresponding theme and the degree measure of existing knowledge j , a knowledge-microblog distribution e w ; The third level, for each word k, corresponding to multiple topics, generating a topic-word distribution s; The generation process of the personalized profiling probability model of Internet users' knowledge and interests is as follows: Process 1, Sampling Distribution: G v,x -Dirichlet (B), L v,w -Dirichlet (T), F v,j -Dirichlet (E), R v,w -Dirichlet (S); Process two, for each user v = 1, 2,..., V, Sampling distribution: H v -Dirichlet(A v ), H-Binomial(π); Sampling distribution: j v -Dirichlet(F v ); Sampling distribution: w v -Dirichlet(R v ); Process three, for each Weibo q = 1, 2,..., Q, Sampling topic: x-Multinomial(H v ) Sampling the depth type z-Multinomial(w of the Weibo description topic v ); Process four, for each word k = 1, 2,..., K, Sampling word: k n,m -Multinomial(G v,x ); Description of the parameters of the user knowledge and interest profiling model: A v , B, T, E, S are hyperparameters and Dirichlet prior distributions; π is a hyperparameter and Binomial prior distribution; H is a |V| * |X| dimensional matrix representing the user-topic distribution; G is a |V| * |X| * |K| dimensional matrix representing the user-topic-word distribution; R is a |V| * |X| * |W| dimensional matrix representing the user-topic-knowledge distribution; F is a |V| * |X| * |J| dimensional matrix representing the user-topic-interest distribution; L is a |V| * |W| * |Z| dimensional matrix representing the user-knowledge-hierarchy distribution; v, j, x, w, z, k are variable instances: user v, user interest j, topic x, existing knowledge w, knowledge level z, word k; V, J, X, W, Z, K represent the set of users, the set of interests, the set of topics, the set of knowledge, the set of hierarchies, and the set of words respectively; Knowledge and interest calculation profiling model: The analysis data input into the model is the user's Weibo data. At the same time, each user corresponds to the user's own Weibo text set. The Weibo texts of M users are used as the input data set of the model. The results have two parts. One part is the user's interest, and the other part is the user's knowledge. In terms of user interest, the model obtains the vector matrix of each user's Weibo text and interest, and calculates the weight of each user for each interest field. The calculation formula is: where x v,j represents the weight of user v for the j-th interest, N represents the number of microblog text sets of user v in the text set, and k v,w,j represents the weight of the j-th interest in the w-th microblog text of user v. For N users, their corresponding interest weights are calculated, and the interest vector of each user is represented as {x v,1 , x v,2 , x v,3 ,…, x v,j ,…}; In terms of user knowledge, all the user's Weibo texts are divided into four categories by the model: super root, super topic, sub topic, and word. The super root represents the root of all super topics. The super topics include the interest-level fields of the user. The sub topics are the link between the super topics and the words, and are the results of the subdivision of the super topics. The bottom-level words are segmented from the user's Weibo. In the model, the number of super topics is set to 1, the number of sub topics is set to 10, the number of words under each sub topic is 5, and there are 50 words at the bottom layer of this hierarchy. The user knowledge hierarchy diagram is generated based on all the data sets and is not specific to a particular user. On the basis of obtaining the user's interest distribution, the words in the user's Weibo are combined with the knowledge level distribution to obtain the description of the user's knowledge level. The depth of the user v's knowledge in a specific interest field is calculated specifically by the following formula: dep(w) v represents the depth of user v's knowledge, and Z represents the depth of the knowledge hierarchy. According to the above model, Z = 4; weight j represents the weight of the node at the j-th layer, and the weight is assigned as: weight j =e j−1 ; N represents the total number of all microblogs of the user; f(j) represents the number of words or topics at the j-th layer that appear in the microblog text set of user v. For repeated words or topics, they are only counted once; show j,i represents whether the word or topic at the j-th layer in the hierarchical structure appears in the microblog text set of user v. If it appears, the value is 1; if not, the value is 0.