Anomaly User Detection Method in Social Networks Based on DS Evidence Theory Fusion
By using DS evidence theory in online social networks to fuse multiple classifiers, combining convolutional neural networks and K nearest neighbor algorithms, abnormal users are detected, and the problem of detection imbalance in the existing technology is solved, and higher detection accuracy and stability are achieved.
Patent Information
- Application Number
- CN202210118942.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-02-08
AI Technical Summary
In the prior art, a single classifier is used to detect abnormal users on online social networks, resulting in unbalanced detection and low accuracy.
A multi-classifier fusion method based on DS evidence theory is used, and a convolutional neural network and K nearest neighbor algorithm is used to detect abnormal users of social networks. By constructing basic probability functions and reliability rules, accurate identification of abnormal users can be achieved.
It improves the accuracy and stability of abnormal user detection, effectively solves the problem of unbalanced detection, and realizes effective detection of abnormal Weibo users.
Smart Images

Figure CN114529762B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network security detection, and in particular to a method for detecting abnormal users in a social network based on DS evidence theory fusion. Background Art
[0002] The number of accounts that publish false or inaccurate messages in the online social network continues to grow, and the huge user information data and user diversity in the platform increase the difficulty of the abnormal user detection problem. How to accurately detect abnormal users from the information dissemination of the online social platform and thus conduct targeted analysis on the abnormal user group is a very meaningful research.
[0003] In the face of a complex network environment, first analyzing the characteristics of abnormal users and detecting abnormal users from the huge user data and published information in the online social network is the basis for the abnormal user detection problem and prevention.
[0004] Currently, the methods for detecting abnormal users in the online social network in the prior art mainly include:
[0005] 1. For the behavioral characteristics of abnormal users, such as the frequency of sending messages or sending a large number of friend requests in a short time, a classifier is used to train these characteristics to construct a detection model.
[0006] 2. Utilizing the fact that the content published by abnormal users is quite different from that of normal users, a classifier is used to train these characteristics to construct a monitoring model.
[0007] The methods for detecting abnormal users using classification models in the above prior art have the following disadvantages: Using a single classifier to detect abnormal users will lead to the problem of unbalanced detection and result in a low detection accuracy rate. Summary of the Invention
[0008] An embodiment of the present invention provides a method for detecting abnormal users in a social network based on DS evidence theory fusion to effectively detect abnormal users on Weibo.
[0009] To achieve the above object, the present invention adopts the following technical solutions.
[0010] A method for detecting abnormal users in a social network based on DS evidence theory fusion includes:
[0011] Constructing and training a convolutional neural network classification model and a K-nearest neighbor algorithm classification model to obtain the accuracy rates of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for detecting abnormal users;
[0012] Use the convolutional neural network classification model and the K-nearest neighbor algorithm classification model respectively to identify the blog text of the detected user, and obtain the detection results of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for the detected user;
[0013] Based on the accuracy rates of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for abnormal user detection through the D-S fusion rule, fuse the detection results of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for the detected user, and obtain the abnormal user detection result of the detected user.
[0014] Preferably, the method further includes:
[0015] Obtain the blog text data published by users on a certain number of online social network platforms, clean and deduplicate the blog text data, remove the emoticons and special symbols in the blog content, perform Chinese word segmentation on the blog text content through the Jieba method, remove the stop words, and obtain the feature vectors of the blog text, which are represented in matrix form;
[0016] Construct a training set and a test set according to the feature vectors of all blog texts.
[0017] Preferably, the constructing and training the convolutional neural network classification model and the K-nearest neighbor algorithm classification model to obtain the accuracy rates of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for abnormal user detection includes:
[0018] Construct an abnormal user classifier model based on the convolutional neural network and an abnormal user classifier model based on the K-nearest neighbor algorithm;
[0019] Use the training set data to train the convolutional neural network classification model and the K-nearest neighbor algorithm classification model, use the test set data to test the convolutional neural network classification model and the K-nearest neighbor algorithm classification model, and obtain the trained abnormal user classifier model based on the convolutional neural network and the abnormal user classifier model based on the K-nearest neighbor algorithm, as well as the average recognition accuracy rates of the two abnormal user classifier models.
[0020] Preferably, the respectively using the convolutional neural network classification model and the K-nearest neighbor algorithm classification model to identify the blog text of the detected user and obtaining the detection results of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for the detected user includes:
[0021] Obtain the blog post text feature vector of the detected user in matrix form, and input this blog post text feature vector into the abnormal user classification model based on the convolutional neural network and the abnormal user classification model based on the K-nearest neighbor algorithm;
[0022] The abnormal user classification model based on the convolutional neural network vectorizes a certain number of blog post text contents of the detected user, uses the learning and training of the hidden layer of the convolutional neural network to mine the deep features of the text, and determines the category detection result of the user to be detected. This category detection result includes the basic probability assignment BPA function, and this BPA function includes whether it is an abnormal user or not;
[0023] The abnormal user classification model based on the K-nearest neighbor algorithm represents the blog post text content in a vector space, classifies the user of a specific category, calculates the similarity between the blog post content of this user and all blog post contents in the training set, then sorts the calculation results in descending order, selects several most similar blog posts, and determines the category detection result of the user to be detected according to the user categories to which these blog posts belong. This category detection result includes the BPA function.
[0024] Preferably, based on the accuracy rate of detecting abnormal users by the D-S fusion rule for the convolutional neural network classification model and the K-nearest neighbor algorithm classification model, fuse the detection results of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for the detected user to obtain the abnormal user detection result of the detected user, including:
[0025] Based on the accuracy rate of detecting abnormal users by the D-S fusion rule for the convolutional neural network classification model and the K-nearest neighbor algorithm classification model, fuse the BPA functions of the detected user on the convolutional neural network classification model and the K-nearest neighbor algorithm classification model to obtain the joint credibility of the detection results of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model, and use the credibility rule according to the joint credibility to obtain the abnormal user detection result of the detected user;
[0026] Let F i (i = 1, 2) respectively represent the convolutional neural network classification model and the K-nearest neighbor algorithm classification model. Input the blog post text feature vector of the detected user into the two classification models, and the obtained recognition result is R i (R i = 0 or R i = 1). When R i = 1, it means the recognition result is an abnormal user, and when R i = 0, it means the recognition result is not an abnormal user. The detection accuracy rate of the i-th classification model for abnormality is P i ;
[0027] The support degree of the abnormal user detection result of the i - type classifier model is initially obtained through the total probability theory formula:
[0028] m i = P i ×R i +(1 - P i )×(1 - R i )
[0029] According to the characteristic that the sum of the belief degrees of the two classifier models on the power set of the recognition framework in the BPA function is equal to 1:
[0030]
[0031] Normalize the above formula to obtain the formula:
[0032]
[0033] where P and R are the recognition accuracy rate and the recognition result respectively;
[0034] According to the above formula, the combined belief degree of the detection results of the convolutional neural network classification model and the K - nearest neighbor algorithm classification model is obtained. According to the DS evidence theory fusion rule and the belief rule, the abnormal user recognition result of the detected user is obtained.
[0035] Let the combined belief degree that the finally detected user is an abnormal user be l(Abn), then l(Abn) should satisfy the following belief rules.
[0036] (1) l(Abn) is the maximum value of the combined belief degree values of the two user attributes.
[0037] (2) The value of l(Abn) must be greater than the threshold x.
[0038] (3) The difference between the objective function l(Abn) and the basic probability assignment value of another category of users must always be greater than the threshold y.
[0039] (4) If the above conditions cannot be satisfied, the output user detection result is "user cannot be recognized".
[0040] It can be seen from the technical solutions provided by the embodiments of the present invention above that the solution of the present invention constructs a basic probability function by combining the recognition results of the detected content on each classifier and the classification accuracy rates of each classifier for different users, and recognizes the detected user after fusing the classifiers through the DS evidence theory fusion rule, realizing the detection of abnormal Weibo users in a balanced and effective manner.
[0041] Additional aspects and advantages of the present invention will be given in part in the following description, which will become apparent from the following description or be understood through the practice of the present invention. Brief Description of the Drawings
[0042] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0043] Figure 1 It is a schematic diagram of the implementation principle of a method for detecting abnormal users in a social network based on the fusion of multi-classifier DS evidence theory provided by the embodiments of the present invention.
[0044] Figure 2 It is a processing flowchart of a method for detecting abnormal users in a social network based on the fusion of multi-classifier DS evidence theory provided by the embodiments of the present invention. Detailed Embodiments
[0045] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation of the present invention.
[0046] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any unit and all combinations of one or more related listed items.
[0047] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as such here.
[0048] For the convenience of understanding the embodiments of the present invention, the following will further explain and illustrate with several specific embodiments in conjunction with the accompanying drawings, and each embodiment does not constitute a limitation to the embodiments of the present invention.
[0049] The present invention proposes a social network abnormal user detection method that combines multiple classifiers and can improve the detection accuracy and stability, that is, the DS evidence theory is used to fuse different classifiers to achieve the detection of abnormal users. The processing process of the method of the present invention includes: using the blog text of users on the online social platform as input, and then mapping and representing the original user text data as feature vectors through data preprocessing. Next, the key representation features of the feature vectors are extracted through the sentence vector model PV-DM. Model training is performed through a convolutional neural network and the K-nearest neighbor algorithm. According to the above two classifier models, the sample set is tested to obtain the accuracy of each classifier in detecting abnormal users. The above two classification models are respectively used to identify the detected users, and the detection results are fused with the average recognition accuracy of the two classifiers to obtain the basic probability function of each classifier for abnormal users. In the DS fusion system, the belief fusion of multiple classifiers for abnormal users is performed to obtain the joint belief that the user to be detected is an abnormal user. Finally, according to the belief rule, the user to be detected is identified to generate the final result.
[0050] The implementation principle diagram of a social network abnormal user detection method based on the DS evidence theory fusion of multiple classifiers provided by the embodiments of the present invention is as Figure 1 shown, and the specific processing flow is as Figure 2 shown, including the following processing steps:
[0051] Step S10: Use the blog text of users on the online social platform as input, and then map and represent the original user text data as feature vectors through the sentence vector model.
[0052] The blog text published by users on the online social network platform may include various forms and expressions. For example, it may include platform emojis, special symbols, and URL links, etc. Obtain a certain amount of blog text data to construct a training set and a test set.
[0053] The input of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model based on text analysis is the matrix representation form of text vectorization. Before inputting the blog text, it is necessary to preprocess the blog text to represent the blog text in a digital form that can be processed by the model. First, clean and deduplicate the blog text data, and remove emojis and special symbols in the blog content, etc. Then, perform Chinese word segmentation on the blog text content through the Jieba method, remove stop words, and obtain the feature vectors of the blog text, which are represented in matrix form.
[0054] The Jieba method scans the word graph based on the prefix dictionary, finds the directed acyclic graph formed by all the generated words of Chinese characters in the sentence, then finds the maximum probability path through dynamic programming to find the maximum segmentation combination based on word frequency. Then, each blog post of the Weibo user is mapped into a paragraph vector through the sentence vector model PV-DM. The paragraph vector is added to the input layer, and each time the paragraph vector participates in the training. And as several words are taken for training by sliding in a sentence, the main idea that the paragraph vector can express will become clearer and more accurate. The 100 Weibo texts of each user are respectively represented by paragraph vectors, and the blog post text content of the user is processed through the sentence vector model PV-DM. The output vector dimension is set to 100, the window size is set to 4, and the number of training iteration rounds is set to 150. The user blog post matrix is represented as a feature vector with a length of 100.
[0055] Step S20: Construct and train a convolutional neural network classification model and a K-nearest neighbor algorithm classification model. According to the above two classifier models, the sample set is tested to obtain the accuracy of each classifier for detecting abnormal users.
[0056] Construct an abnormal user classifier model based on a convolutional neural network and an abnormal user classifier model based on the K-nearest neighbor algorithm.
[0057] Use the training set data to train the above two classifier models, and use the test set data to test the above two classifier models to obtain the trained abnormal user classifier model based on a convolutional neural network and the abnormal user classifier model based on the K-nearest neighbor algorithm, as well as the average recognition accuracy of the two abnormal user classifier models.
[0058] Step S30: Respectively use the above two classifier models to identify the detected users, and fuse the detection results and the average recognition accuracy of the two classifiers to obtain the basic probability function of each classifier for abnormal users.
[0059] Obtain the blog post text feature vector of the detected user represented in matrix form, and input the blog post text feature vector into the above abnormal user classifier model based on a convolutional neural network and the abnormal user classifier model based on the K-nearest neighbor algorithm.
[0060] The anomaly user classifier model based on a convolutional neural network vectorizes a certain number of blog text contents of the user to be detected, and uses the learning and training of the hidden layer of the convolutional neural network to mine the deep features of the text, and determines the category detection result of the user to be detected. The category detection result includes a Basic Probability Assignment (BPA) function, and the BPA function includes whether the user is an abnormal user or not. In this way, manual feature construction is avoided, and abnormal users can be identified even when the user information obtained is insufficient.
[0061] The anomaly user classifier model based on the K-nearest neighbor algorithm represents the blog text content in a vector space, classifies the user of a specific category, calculates the similarity between the blog content of this user and all blog contents in the training set, then sorts the calculation results in descending order, selects several most similar blogs, and determines the category detection result of the user to be detected according to the user categories to which these blogs belong. The category detection result includes a BPA function, and the BPA function includes whether the user is an abnormal user or not.
[0062] Then, based on the average recognition accuracy of the above two anomaly user classifier models through the Dempster-Shafer (D-S) fusion rule, the BPA functions of the user to be detected on the above two anomaly user classifier models are fused to obtain the combined belief of the detection results of the above two anomaly user classifier models for the user to be detected. According to the combined belief, the anomaly user detection of the above user to be detected is obtained using the belief rule.
[0063] In one embodiment, let F i (i = 1, 2) respectively represent the anomaly user classifier model based on a convolutional neural network and the anomaly user classifier model based on the K-nearest neighbor algorithm. Input the blog content feature vector of the online social network user into the two anomaly user classifier models, and obtain the detection accuracy of the i-th type of anomaly user classifier model for anomalies as P i .
[0064] Introduce the blog content of the user to be detected, and perform recognition on the two anomaly user classifier models respectively. The obtained recognition results are R i (R i = 0 or R i = 1). When R i = 1, it means the recognition result is an abnormal user, and when R i = 0, it means the recognition result is not an abnormal user. Then, through the total probability theory formula, the support degree of the anomaly user detection result of the i-th type of anomaly user classifier model is initially obtained:
[0065] m i = P i×R i +(1 - P i )×(1 - R i )
[0066] According to the feature that the sum of the belief degrees of two classifier models on the power set of the recognition framework in the BPA function is equal to 1:
[0067]
[0068] Normalize the above formula to obtain the formula:
[0069]
[0070] where P and R are the recognition accuracy rate and recognition result respectively.
[0071] According to the above formula, the belief degree values of the two abnormal user classifier models for abnormal users can be obtained, and the recognition result of the detected user can be obtained according to the DS evidence theory fusion rule and belief rule.
[0072] Let the combined belief degree that the finally detected user is an abnormal user be l(Abn), then l(Abn) should satisfy the following belief rules.
[0073] (1) l(Abn) is the maximum value of the combined belief degree values of the two user attributes.
[0074] (2) The value of l(Abn) must be greater than the threshold x.
[0075] (3) The difference between the objective function l(Abn) and the basic probability assignment value of another category of users must always be greater than the threshold y.
[0076] (4) If the above conditions cannot be satisfied, the output user detection result is "user cannot be recognized".
[0077] According to the belief rules to determine the category of the finally detected user. After experiments, the present invention determines the value of x as 0.80 and the value of y as 0.52.
[0078] In summary, the solution of the present invention realizes the detection of abnormal users in the online social network by fusing classifiers through the DS evidence theory fusion rule, and achieves the detection of abnormal users in an equilibrium and effective manner.
[0079] The method of the embodiment of the present invention has a higher abnormal user detection accuracy rate and higher anti-interference ability compared with the prior art solution.
[0080] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0081] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0082] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0083] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for detecting abnormal users in social networks based on D-S evidence theory fusion, characterized in that, it includes: Construct and train a convolutional neural network classification model and a K-nearest neighbor algorithm classification model to obtain the accuracy rates of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for detecting abnormal users; Use the convolutional neural network classification model and the K-nearest neighbor algorithm classification model respectively to identify the blog text of the detected user, and obtain the detection results of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for the detected user; Based on the accuracy rates of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for detecting abnormal users through the D-S fusion rule, fuse the detection results of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for the detected user to obtain the abnormal user detection result of the detected user, specifically including: Based on the accuracy rates of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for detecting abnormal users through the D-S fusion rule, fuse the BPA functions of the detected user on the convolutional neural network classification model and the K-nearest neighbor algorithm classification model to obtain the joint credibility of the detection results of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model, and obtain the abnormal user detection result of the detected user according to the joint credibility using the credibility rule; Let F i (i = 1, 2) respectively represent the convolutional neural network classification model and the K-nearest neighbor algorithm classification model. Input the blog text feature vector of the user to be detected into the two classifier models, and the recognition result is R i (R i = 0 or R i = 1). When R i = 1, it means the recognition result is an abnormal user. When R i = 0, it means the recognition result is not an abnormal user. The detection accuracy rate of the i-class classifier model for abnormality is P i ; Obtain the support degree of the abnormal user detection result of the i-th classifier model preliminarily through the total probability theory formula: m i = P i × R i + (1 - P i ) × (1 - R i ) According to the characteristic that the sum of the credibilities of the two classifier models on the power set of the recognition framework of the BPA function is equal to 1: Normalize the above formula to obtain the formula: where P and R are the recognition accuracy rate and the recognition result respectively; Obtain the joint credibility of the detection results of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model according to the above formula, and obtain the abnormal user recognition result of the detected user according to the D-S evidence theory fusion rule and the credibility rule; Let the joint credibility that the finally detected user is an abnormal user be l(Abn), then l(Abn) should satisfy the following credibility rules; (1) l(Abn) is the maximum value of the joint credibility values of the two user attributes; (2) The value of l(Abn) must be greater than the threshold x; (3) The difference between the objective function l(Abn) and the basic probability assignment value of another category of users must always be greater than the threshold y; (4) If the above conditions cannot be satisfied, the output user detection result is "user cannot be recognized".
2. The method according to claim 1, characterized in that, the method further includes: Obtain the blog text data published by users in a certain number of online social network platforms, clean and deduplicate the blog text data, remove the emoticons and special symbols in the blog content, perform Chinese word segmentation on the blog text content through the Jieba method, and remove the stop words to obtain the feature vector of the blog text, and this feature vector is represented in matrix form; Construct a training set and a test set according to the feature vectors of all blog texts.
3. The method according to claim 2, characterized in that, Building and training a convolutional neural network classification model and a K-nearest neighbor algorithm classification model to obtain the accuracy rates of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for detecting abnormal users, including: Building an abnormal user classifier model based on a convolutional neural network and an abnormal user classifier model based on the K-nearest neighbor algorithm; Training the convolutional neural network classification model and the K-nearest neighbor algorithm classification model using the training set data, and testing the convolutional neural network classification model and the K-nearest neighbor algorithm classification model using the test set data to obtain a trained abnormal user classifier model based on the convolutional neural network and an abnormal user classifier model based on the K-nearest neighbor algorithm, as well as the average recognition accuracy rates of the two abnormal user classifier models.
4. The method according to claim 3, wherein, Respectively using the convolutional neural network classification model and the K-nearest neighbor algorithm classification model to identify the blog text of the user to be detected, and obtaining the detection results of the convolutional neural network classification model and the K-nearest neighbor algorithm classification model for the user to be detected, including: Obtaining the blog text feature vector of the user to be detected represented in matrix form, and inputting this blog text feature vector into the abnormal user classifier model based on the convolutional neural network and the abnormal user classifier model based on the K-nearest neighbor algorithm; The abnormal user classifier model based on the convolutional neural network vectorizes a certain number of blog text contents of the user to be detected, and uses the learning and training of the hidden layer of the convolutional neural network to mine the deep features of the text, and determines the category detection result of the user to be detected. This category detection result includes a basic probability assignment BPA function, and this BPA function includes whether the user is an abnormal user or not; The abnormal user classifier model based on the K-nearest neighbor algorithm represents the blog text content in a vector space, classifies the user of a specific category, calculates the similarity between the blog content of this user and all the blog contents in the training set, then sorts the calculation results in descending order, selects several most similar blogs, and determines the category detection result of the user to be detected according to the user categories to which these blogs belong. This category detection result includes a BPA function.