Network virtual identity behavior prediction analysis method and system
By introducing distributed computing, stream processing technology and deep learning algorithms in user behavior data analysis, and combining BERT model for text similarity analysis, the problems of high computational complexity, insufficient analysis depth and poor real-time performance in the existing technology are solved, and efficient and accurate user behavior prediction and personalized recommendation are achieved.
Patent Information
- Application Number
- CN202411935793.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
AI Technical Summary
When processing massive and high-dimensional user behavior data, the prior art has high computational complexity, insufficient analysis depth, poor real-time performance, and difficult to achieve accurate user behavior prediction and personalized recommendation.
Distributed computing and stream processing technology combined with deep learning algorithms are used to analyze text similarity through BERT model, calculate file, event and period similarity, and predict user behavior trends and potential risks through regression prediction calculations.
It improves the depth and real-time nature of data analysis, achieves more accurate user behavior prediction and personalized recommendation, improves analysis efficiency and accuracy, and supports security monitoring and decision-making.
Smart Images

Figure CN120068012A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method and system for predicting and analyzing network virtual identity behaviors. Background Art
[0002] Analysis of network virtual user behaviors is a key technology that involves collecting, processing, analyzing, and interpreting user behavior data in the cyber space to reveal user behavior patterns, preferences, needs, and potential risks. Current analysis of network virtual user behaviors mainly focuses on in-depth analysis of user behavior patterns, characteristics, and interaction data on the network. Generally, it mainly includes the following specific processes:
[0003] 1. Data collection: Collect user behavior data, including browsing history: including visited websites, page stay times, clicked links, etc., which can reflect user interests and areas of concern; interaction behaviors: such as likes, comments, and shares on social media, posting and replying in forums, etc., which can reflect user opinion tendencies and social interaction patterns; search records: user search keywords can reveal their current needs and concerns; collect network environment data, including time and location information: different times and locations may affect user network behaviors. For example, user behaviors may vary between working hours and at home; device information: the type of device used (such as mobile phone, computer, tablet), operating system, etc. may also be related to specific behavior patterns.
[0004] 2. Feature extraction: including activity behavior features, interest preference features, social features, etc.
[0005] 3. Model selection: generally including machine learning methods (classification algorithms, regression algorithms, clustering algorithms), deep learning methods (recurrent neural network (RNN), long short-term memory network (LSTM), convolutional neural network (CNN)), etc.
[0006] 4. Evaluation: Comprehensive evaluation is carried out through evaluation metrics (accuracy rate which is the proportion of correctly predicted behaviors in the total predicted behaviors, recall rate which is the proportion of actually occurring behaviors that are correctly predicted, F1 value which comprehensively considers the accuracy rate and recall rate, and prediction error used to evaluate the regression model).
[0007] For example, in the prior art, a patent for invention with a publication number of CN115345390A discloses a method, apparatus, electronic device, and storage medium for predicting behavioral trajectories. By collecting multi-dimensional trajectory data, it solves the problem of incomplete trajectory coverage in the prior art solutions. At the same time, by determining the relationship association between the multi-dimensional trajectory data and the true identity of the target person, the data in the target trajectory database can all be associated with the true identity of the target person, improving the efficiency of determining the identity of the target person through various trajectory information. Finally, by establishing computational models for different behavioral feature classifications to respectively perform statistical analysis on features of different dimensions, and establishing a machine learning model, and continuously updating the weights of the computational models for different behavioral feature classifications according to the prediction results of the machine learning model, the prediction accuracy of the finally obtained comprehensive trajectory analysis prediction model is further improved. Compared with a single statistical computational model or machine learning model that cannot take into account both the accuracy of the prediction result and the relevance to the original data, the behavioral trajectory prediction method provided by the present invention combines the two, and further improves the accuracy of the prediction result on the premise of ensuring the relevance between the prediction result and the original data.
[0008] As can be seen from the above, the current data analysis methods rely on the selection of machine learning models. Existing models, whether they are single statistical computational models or machine learning models, or a combination of multiple models, often have problems such as high computational complexity, insufficient analysis depth, and poor real-time performance when dealing with massive and high-dimensional user behavior data. Summary of the Invention
[0009] A brief overview of the embodiments of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that the following overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to a more detailed description to follow.
[0010] To solve the deficiencies of the prior art, the present application introduces computational methods of distributed computing, stream processing technology, and deep learning algorithms. On the one hand, running the deep learning algorithm on the distributed computing framework and stream processing technology can accelerate the timeliness and speed of the algorithm. On the other hand, the stream processing technology is based on distributed computing, and the deep learning algorithm is applied on both of them, thereby improving the depth and real-time performance of data analysis and achieving more accurate user behavior prediction and personalized recommendation. The present application comprehensively and deeply analyzes the network virtual user behavior in an intelligent and automated manner, not only improving the analysis efficiency and accuracy, tracking the changes in user behavior, and providing strong support for decision-making.
[0011] This application conducts text similarity analysis through the BERT model. Compared with existing analysis models, the data set selected in this application is the titles and keywords extracted from text files related to user behavior. And in the specific implementation, the streaming processing technology is used to solve the problems of poor real-time performance, high computational complexity, and large data volume. To improve the depth of data analysis, the overall calculation is split into small parallel calculations, thus solving the problem of high computational complexity.
[0012] According to one aspect of the present application, there is provided a method for predicting and analyzing network virtual identity behavior, which includes:
[0013] Calculating the file similarity, event similarity, and cycle similarity processed by the user to be analyzed through the BERT model; compared with existing analysis models, the data set selected in this application is the titles and keywords extracted from text files related to user behavior, and in the specific implementation, the event similarity and cycle similarity processed by the user to be analyzed can be more reliable and advanced, thus solving the problem of high computational complexity.
[0014] Aggregating the file similarity, event similarity, and cycle similarity processed by the user to be analyzed by virtual identity classification;
[0015] Performing regression prediction calculations on the virtual identity data after classification aggregation to predict the behavior trends and potential risks of each virtual identity;
[0016] According to the results of the regression prediction, select virtual identities with higher risks to obtain a topN result set; the topN result set represents the virtual identities most likely to have dangerous behaviors. In addition, the value of N can be adjusted according to specific requirements and scenarios to obtain more accurate results;
[0017] Supplement the location information of the virtual identity; its location information can be obtained through methods such as IP addresses and GPS positioning;
[0018] Finally, obtain the virtual identity personnel with dangerous behaviors.
[0019] As a specific implementation, calculating the file similarity processed by the user to be analyzed according to the BERT model specifically includes:
[0020] Step 1, data collection and preprocessing: Collect a set of text files related to user behavior, and extract the title and keywords from each text file in the text file set as the preprocessed text; among them, the title and keywords can be accurately extracted through specific text analysis algorithms and tools.
[0021] Step 2: Read one text file from the set of text files as the current text file, and use the BERT model to compare and analyze the similarity between the title and keywords of the current text file and the keywords of historical files in the database. During the comparison process, the similarity between the title and each keyword in the current text file and all keywords in the historical files will be calculated one by one.
[0022] Step 3: For each group of keywords in the current text file, after comparing with the keywords of the historical files, select the maximum similarity value. This maximum value represents the degree of closeness between this group of keywords and the most similar keywords in the historical files. In this way, the strongest association between the keywords of the current file and the keywords of the historical files can be found.
[0023] Step 4: Calculate the file similarity value: Take the average of multiple maximum values as the similarity value of the current text file. However, if the similarity value is 0, it will not be included in the calculation. A similarity value of 0 means there is no similarity. Therefore, only when there is a certain degree of similarity will the historical file be considered. This can ensure that the final similarity value obtained is meaningful and can accurately reflect the similarity degree between files.
[0024] Step 5: Bind the historical files to each user (node information). At the same time, exclude the data of the current user (node) during the analysis process, and extract the similarity value and average value data of each user (node)'s files.
[0025] Step 6: Repeat Steps 2 - 5 to obtain the similarity value and average value data of all text files in the set of text files as the file similarity.
[0026] Further, in Step 4, the similarity calculation formula for each file is:
[0027] cosθ = (A·B) / (||A|| × ||B||);
[0028] In the formula, cosθ is the cosine similarity value, and θ represents the angle between vector A (event) and vector B (event); A·B represents the dot product of vector A (event) and vector B (event), and ||A|| and ||B|| are the norms (i.e., the lengths of the vectors) of vector A (event) and vector B (event) respectively, which can be obtained by taking the square root of the sum of the squares of each element.
[0029] The value range of the cosine similarity cosθ is between -1 and 1. When the directions of two vectors are exactly the same (the included angle is 0°), the cosine similarity is 1; when the directions of two vectors are exactly opposite (the included angle is 180°), the cosine similarity is -1; when two vectors are orthogonal (the included angle is 90° or 270°), the cosine similarity is 0. Therefore, the closer the value of the cosine similarity is to 1, the more consistent the directions of the two vectors are, and the more similar they are; conversely, the closer the value is to -1, the more opposite the directions of the two vectors are, and the less similar they are. Events with the same similarity belong to the same category.
[0030] Specifically, calculating the similarity of events processed by the analyzed user according to the BERT model specifically includes:
[0031] Step 11: Describe the event in text form, extract the text content related to the event, and then remove the noise in the text content to obtain the cleaned text; extract valuable feature words from the cleaned text to form a preprocessed text file; the valuable feature words are the feature words that measure their importance in the text content, including that their frequency of occurrence in a single text file is greater than a preset value, and their rarity in the entire text collection is within a preset threshold range;
[0032] Step 12: Read one text file from the set of preprocessed text files as the current text file, and compare and analyze the similarity between the feature words of the current text file and the feature words of the historical files in the database through the BERT model; during the comparison process, the similarity between the feature words in the current text file and all the feature words in the historical files will be calculated one by one;
[0033] Step 13: For each group of feature words in the current text file, after comparing with the feature words of the historical file, select the maximum similarity value; this maximum value represents the closeness between this group of feature words and the most similar feature words in the historical file. In this way, the strongest association between the feature words of the current file and the feature words of the historical file can be found;
[0034] Step 14: Calculate the file similarity value: Take the average of multiple maximum values as the similarity value of the current text file; however, if the similarity value is 0, it will not be included in the calculation;
[0035] Step 15: Bind the historical files to each user. At the same time, exclude the data of the current user during the analysis process, and extract the similarity value and average value data of each user's file;
[0036] Step 16: Repeat Step 12 - Step 15 to obtain the similarity value and average value data of all text files in the text file set as the event similarity.
[0037] Specifically, calculating the period similarity processed by the user to be analyzed according to the BERT model specifically includes:
[0038] Step 21, time series data extraction: For periodic events, extract time series data. At the same time, fill in the missing data using interpolation methods to ensure the accuracy and integrity of the time series data; Periodic feature extraction: Determine the period length of the data, and extract periodic features according to the period length of the data; The periodic features form a preprocessed text file;
[0039] Step 22, read one of the text files in the preprocessed text file set as the current text file, and compare and analyze the similarity between the periodic features of the current text file and the features of the historical files in the database through the BERT model; During the comparison process, the similarity between the periodic features in the current text file and all features in the historical files will be calculated one by one;
[0040] Step 23, for each group of periodic features of the current text file, after comparing with the historical file features, select the maximum similarity value; This maximum value represents the degree of closeness between this group of periodic features and the most similar features in the historical file. In this way, the strongest association between the current file features and the historical file features can be found;
[0041] Step 24, calculate the file similarity value: Take the average of multiple maximum values as the similarity value of the current text file; However, if the similarity value is 0, it will not participate in the calculation;
[0042] Step 25, bind the historical files to each user, and at the same time, exclude the current user data during the analysis process, and extract the similarity value and average value data of each user's file;
[0043] Step 26, repeat Step 22 - Step 25 to obtain the similarity value and average value data of all text files in the text file set as the period similarity.
[0044] Among them, binding the historical files to each user (node information) is used to obtain the events corresponding to the files related to or similar to the current file. Through this binding, it is more convenient to trace and understand the association between similar files, providing more background information for further analysis and decision-making. Excluding the data of the current node during the analysis process to avoid repeated calculations and interference, so as to ensure that only the similarity between the historical file and the current file is concerned, in order to more accurately extract valuable information. Extracting the similarity and average value data of each node's file can be used for various purposes such as recommending relevant files, conducting trend analysis, and discovering potential associations.
[0045] Further, obtain the approval processing information of the tasks sent by the current node user (i.e., the data reviewed after algorithm calculation). According to the similarity of the key information and the processing information in the historical data, exclude the current user's data. Take the maximum similarity value of each group of keywords in the current approval and the keywords in the historical files, and then take the average value of multiple maximum values as the similarity value of the approval. If the similarity value is 0, it does not participate in the calculation. Bind the historical data to each node information and take the average value of each node.
[0046] As a specific implementation plan, in addition to the similarity of the files processed by the analyzed user, the similarity of the events processed by the analyzed user, and the similarity of the cycles processed by the analyzed user as the main calculation indicators, it also includes whether the analyzed user has completed the event, the current amount of events completed by the analyzed user, and the ratio of the completed amount to the access amount of the analyzed user as auxiliary calculation indicators.
[0047] As a specific implementation plan, the calculation rules of the auxiliary calculation indicators are as follows:
[0048] Obtain whether the current user is online. If online, add 1 to the score;
[0049] Obtain the events that the current user has accessed. If there are multiple events, then add 1 to each event. Events with the same similarity are of one category;
[0050] Query the number of events currently completed by the user. If there are multiple completed events, then take out all the completed events. Events with the same similarity are of one category;
[0051] Completion access ratio (T) = completed events (ST) / sum of accessed events (AT);
[0052] Filter out the cases where the completion access ratio (T) is 1, the completed events (ST) is 1, and the sum of accessed events (AT) is 1;
[0053] If the completion access ratio (T) is higher, it means that the user is more real and the prediction is more accurate.
[0054] As a specific implementation plan, the detailed calculation method of this prediction analysis method is as follows:
[0055] Events with the same similarity are of one category:
[0056]
[0057] Calculate the value at A. A depends on ra and pa. ra represents the cosine similarity of the completed events (ST), and pa represents the completion access ratio (T);
[0058] Prediction algorithm calculation formula:
[0059]
[0060] Let x represent an event that has been completed, and y represent the events viewed each day; X is the set of events x completed each day, Y is the set of events viewed each day, n is the number of days, x n represents the event completed on the nth day, and y n represents the event viewed on the nth day;
[0061] First, it is necessary to calculate the average values of X and Y. x / n is the mean of X; y / n is the mean of Y;
[0062] Calculate the regression coefficient f(x, y): Use the above formula to calculate the regression coefficient f(x, y), which represents the strength of the linear relationship between X and Y;
[0063] Finally, using the obtained regression coefficient f(x, y) and the average values of X and Y, a linear regression equation can be constructed y= b x+ay= b x+a , where a and b are the coefficients of the linear regression equation.
[0064] Based on the above calculations, prediction information such as public opinion publication, real-time comments, and prohibited items for each virtual identity in the n-day cycle is obtained, and a top-N sorted result set is generated. Predictive judgments are made on the subsequent behaviors of virtual identities by taking the top ones from the result set. Geographical locations are supplemented according to the predictive judgment results of each virtual identity to achieve the purpose of this predictive judgment calculation.
[0065] Furthermore, the virtual identity behavior prediction analysis performs real-time calculations on the different events generated by each person, and generates an event sorting form in reverse order within the most recent window; different events generated by each person are extracted from contents such as text, base stations, and logs, including location events, click events, and completion events.
[0066] Because there are many factors involved in the prediction analysis method, if only one reverse-order form is generated, the risk of failure is relatively high, and it is also difficult to maximize the benefits of the entire prediction analysis. As a further solution, this application takes the union of the reverse-order forms generated for the past time window (event time), checks the proportion of the current form in the union, and sorts and selects values according to the proportion.
[0067] As a specific implementation, this prediction analysis method includes individual subsequent behavior prediction and group subsequent behavior prediction, and sorts and selects values according to the proportion of the above results as future result predictions.
[0068] As a specific implementation, before data collection and preprocessing, it further includes the steps of pre-training the BERT model and adjusting the pre-trained BERT model with the collected text files, so as to update the parameters of the BERT model and make it adapt to specific user behavior prediction tasks.
[0069] According to another aspect of the present application, there is provided a network virtual identity behavior prediction and analysis system for performing the above prediction and analysis method, which includes:
[0070] A similarity calculation module for calculating the file similarity, event similarity, and cycle similarity processed by the user to be analyzed according to the BERT model;
[0071] A classification and aggregation module for classifying and aggregating the file similarity, event similarity, and cycle similarity processed by the user to be analyzed according to the virtual identity;
[0072] A regression prediction calculation module for performing regression prediction calculations on the virtual identity data after classification and aggregation to predict the behavior trends and potential risks of each virtual identity;
[0073] A result output module for selecting virtual identities with higher risks according to the results of the regression prediction to obtain a topN result set; supplementing location information for the virtual identities in the topN result set to finally obtain virtual identity personnel with dangerous behaviors.
[0074] The present application realizes a virtual identity behavior prediction and analysis method through the above solutions. By mainly calculating indicators (file similarity, event similarity, cycle similarity processed by the user to be analyzed), it conducts network virtual user behavior analysis, realizes comprehensive and in-depth analysis of network virtual user behavior, improves the analysis efficiency and accuracy, tracks user behavior changes, and provides strong support for decision-making; by innovatively binding historical files to each user to obtain events corresponding to files related to or similar to the current file, through this binding, it is more convenient to trace and understand the associations between similar files, providing more background information for further analysis and decision-making; it also innovatively introduces the calculation of event similarity and cycle similarity, successfully capturing the similarity of events and the periodic characteristics in the time series, which can better assist in the prediction of virtual identities.
[0075] Meanwhile, through the collaborative analysis of auxiliary calculation indicators (whether the user to be analyzed has completed an event, the current amount of events completed by the user to be analyzed, and the ratio of the completion amount to the access amount of the user to be analyzed), the depth and real-time performance of data analysis are improved, enabling more accurate prediction of user behavior and personalized recommendations. The predictive analysis method implemented by the present application in an intelligent and automated manner has important application value, can provide strong support for security monitoring, etc., and can be applied to the security guarantee of major events, relevant venues, the analysis of dangerous behaviors of network virtual personnel, and the periodic attention to dangerous behaviors, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The present invention can be better understood by referring to the descriptions given below in conjunction with the accompanying drawings, in which the same or similar reference numerals are used throughout the drawings to denote the same or similar components. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of this specification, and are further used to illustrate the preferred embodiments of the present invention and to explain the principles and advantages of the present invention. In the
[0077] In the figures:
[0078] Figure 1 is a principle block diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0079] Embodiments of the present invention will be described below with reference to the accompanying drawings. Elements and features described in one drawing or one embodiment of the present invention may be combined with elements and features shown in one or more other drawings or embodiments. It should be noted that, for the sake of clarity, the drawings and the description omit the representation and description of components and processes that are irrelevant to the present invention and are known to those of ordinary skill in the art.
[0080] Traditional data analysis methods often have problems such as high computational complexity, insufficient analysis depth, and poor real-time performance when dealing with massive and high-dimensional user behavior data. The present invention introduces advanced algorithms such as distributed computing, stream processing technology, and deep learning to improve the depth and real-time performance of data analysis, and achieve more accurate prediction of user behavior and personalized recommendations.
[0081] The present invention performs text similarity analysis through the BERT model. Compared with existing analysis models, the data set selected in the present application is the titles and keywords extracted from text files related to user behavior, and in the specific implementation, the stream processing technology is used to solve the problems of poor real-time performance and high computational complexity, and the large amount of data improves the depth of data analysis, and the overall calculation is split into small parallel calculations, thus solving the problem of high computational complexity.
[0082] The basic principle of text similarity analysis using BERT is as follows: First, the text to be compared is input into the BERT model. The BERT model will convert the text into a series of word vectors (Token Embeddings) and positional vectors (Positional Embeddings), and through the self-attention mechanism, fuse the word vectors and positional vectors to generate a representation vector (Contextual Representation) for each text. Then, the similarity between the two text representation vectors (such as cosine similarity) is calculated to measure the similarity between the texts.
[0083] Predict individual and group subsequent behaviors through personal intelligence analysis, and use it as a judgment of the behaviors and relationships of individuals and groups.
[0084] This analysis method performs real-time calculations on different events generated by each pair of characters, including location events, click events, completion events, etc., and generates an event sorting form in reverse order within the most recent window.
[0085] Randomness: Since the prediction analysis method involves many factors, if only one reverse-order form is generated, the risk of failure is relatively high, and it is also difficult to maximize the benefits of the entire prediction analysis.
[0086] Coverage rate: Take the union of the reverse-order forms generated for the past time window and see the proportion of the current form in the union.
[0087] Credibility: This indicator is relatively subjective. It means that the choices made for the completion events that users trust are reasonable, and how the prediction analysis method operates internally.
[0088] As a specific embodiment, the present invention provides a method for predicting and analyzing the behaviors of network virtual identities, which includes:
[0089] Calculate the similarity of the files processed by the user to be analyzed through the BERT model. Compared with the existing analysis models, the data set selected in this application is the titles and keywords extracted from text files related to the user's behavior, and in the specific implementation, the similarity of the events processed by the user to be analyzed and the similarity of the cycles processed by the user to be analyzed; it can be more credible and advanced, thus solving the problem of high computational complexity.
[0090] Aggregate the similarity of the files processed by the user to be analyzed, the similarity of the events processed by the user to be analyzed, and the similarity of the cycles processed by the user to be analyzed according to the virtual identity.
[0091] Perform regression prediction calculations on the virtual identity data after classification and aggregation to predict the behavior trends and potential risks of each virtual identity.
[0092] According to the results of regression prediction, select virtual identities with higher risks to obtain a top-N result set; the top-N result set represents virtual identities that are most likely to have dangerous behaviors; in addition, the value of N can be adjusted according to specific requirements and scenarios to obtain more accurate results;
[0093] Supplement the location information of the virtual identity; its location information can be obtained through methods such as IP addresses and GPS positioning;
[0094] Finally, obtain the virtual identity personnel with dangerous behaviors.
[0095] As a specific implementation, calculating the file similarity, event similarity, and cycle similarity of the analyzed user according to the BERT model specifically includes:
[0096] 1. File similarity
[0097] Calculating the file similarity includes converting the file content into vector representations and then comparing these vectors through cosine similarity or other similarity measurement methods. Specifically, it includes:
[0098] Step 1, data collection and preprocessing: Collect a set of text files related to user behaviors, and extract the title and keywords from each text file in the set of text files as the preprocessed text; among them, extracting the title and keywords can accurately extract these key elements through specific text analysis algorithms and tools;
[0099] Step 2, read one text file from the set of text files as the current text file, and compare and analyze the similarity between the title and keywords of the current text file and the keywords of historical files in the database through the BERT model; during the comparison process, the similarity between the title and each keyword in the current text file and all keywords in the historical file will be calculated one by one;
[0100] Step 3, for each group of keywords in the current text file, after comparing with the historical file keywords, select the maximum value of the similarity; this maximum value represents the closeness between this group of keywords and the most similar keywords in the historical file. In this way, the strongest association between the current file keywords and the historical file keywords can be found;
[0101] Step 4, calculate the file similarity value: Take the average of multiple maximum values as the similarity value of the current text file; however, if the similarity value is 0, it will not participate in the calculation. A similarity value of 0 means there is no similarity. Therefore, only when there is a certain degree of similarity will the historical file be included in the consideration range, so that the finally obtained similarity value has practical significance and can accurately reflect the similarity degree between files;
[0102] Step 5: Bind the historical files to each user. Meanwhile, during the analysis process, exclude the current user's data and extract the similarity values and average values of each user's file.
[0103] Step 6: Repeat Steps 2 - 5 to obtain the similarity values (such as cosine similarity) and average values ( where x represents an event that has been completed, n represents the number of days, and x i represents the event completed on the i-th day) of all text files in the text file set as the file similarity.
[0104] 2. Event Similarity
[0105] The calculation of event similarity is achieved by converting the event description into a vector representation and then calculating the similarity between these vectors.
[0106] Step 11: Text extraction and cleaning: If the event is described in text form, first extract the text content related to the event, including the event title, description, comments, etc. Then remove the noise in the text to obtain the cleaned text, such as HTML tags, special symbols, stop words (words like "of", "is", "in", etc. that are not very important for semantic judgment). For example, for a news event report, use tools like regular expressions to clean the text.
[0107] Feature extraction: Extract valuable feature words from the cleaned text as the preprocessed text file. Specifically, the bag-of-words model can be used, considering the text as a set of words and counting the frequency of each word appearance as the feature vector. This method ignores the word order and grammar structure. Another better method can also be used - TF-IDF (Term Frequency-Inverse Document Frequency), which not only considers the frequency of a word in a single text but also takes into account the rarity of the word in the entire text set to measure the importance of the word to the text;
[0108] The other steps are roughly the same as Steps 2 - 6 of the above file similarity and are described as follows:
[0109] Step 12: Read one text file from the preprocessed text file set as the current text file, and compare and analyze the similarity between the feature words of the current text file and the feature words of the historical files in the database through the BERT model; during the comparison process, the similarity between the feature words in the current text file and all feature words in the historical files will be calculated one by one;
[0110] Step 13: For each group of characteristic words in the current text file, after comparing with the characteristic words of the historical file, select the maximum similarity value; this maximum value represents the closeness between this group of characteristic words and the most similar characteristic words in the historical file. In this way, the strongest association between the characteristic words of the current file and the historical file can be found.
[0111] Step 14: Calculate the file similarity value: Take the average of multiple maximum values as the similarity value of the current text file; however, if the similarity value is 0, it will not be involved in the calculation.
[0112] Step 15: Bind the historical file to each user. At the same time, exclude the current user data during the analysis process, and extract the similarity value and average value data of each user's file.
[0113] Step 16: Repeat Steps 12 - 15 to obtain the similarity value and average value data of all text files in the text file set as the event similarity.
[0114] 3. Period similarity:
[0115] The calculation of period similarity includes the processing of time series data, which is achieved by converting time series data into vector representations and calculating the similarity between these vectors.
[0116] Step 21: Time series data extraction: For events with periodicity, extract time series data. For example, for sales data, record the sales amount for each time period (such as daily, weekly, monthly). At the same time, to ensure the accuracy and integrity of time series data, interpolation and other methods can be used to fill in the missing data. Periodic feature extraction: Determine the period length of the data, and extract periodic features according to the period length of the data. Among them, the main period of the data sequence can be found through autocorrelation analysis. For example, for website traffic data with an obvious daily period, by analyzing the autocorrelation function, it can be found that the data shows a similar fluctuation pattern every 24 hours, which is one of its periods. The periodic features are used as the preprocessed text files.
[0117] The other steps are roughly the same as Steps 2 - 6 of file similarity and are described as follows:
[0118] Step 22: Read one text file from the preprocessed text file set as the current text file, and compare and analyze the similarity between the periodic features of the current text file and the features of the historical files in the database through the BERT model; during the comparison process, the similarity between the periodic features in the current text file and all features in the historical file will be calculated one by one.
[0119] Step 23: For each set of periodic features in the current text file, after comparing with the historical file features, select the maximum similarity value; this maximum value represents the tightness between this set of periodic features and the most similar features in the historical file. In this way, the strongest association between the current file features and the historical file features can be found.
[0120] Step 24: Calculate the file similarity value: Take the average of multiple maximum values as the similarity value of the current text file; however, if the similarity value is 0, it does not participate in the calculation.
[0121] Step 25: Bind the historical file to each user. At the same time, exclude the current user data during the analysis process, and extract the similarity value and average value data of each user's file.
[0122] Step 6: Repeat Steps 22 - 25 to obtain the similarity value and average value data of all text files in the text file set as the periodic similarity.
[0123] The present invention innovatively introduces the calculation of event similarity and periodic similarity, which can successfully capture the similarity of events and the periodic features in the time series, and can better assist in the prediction of virtual identities.
[0124] Further, in Step 1 of the periodic similarity, when extracting time series data, in specific implementation, the Fourier transform is introduced for implementation. The Fourier transform (FFT) is a mathematical tool that transforms time series data from the time domain to the frequency domain, and can effectively identify and extract the periodic features in the time series. Through the Fourier transform, the time series can be decomposed into the superposition of different frequency components, making it easier to identify the periodic patterns in the data. Specifically, for periodic events, extract the time series data as the preprocessed text file, and use the fast Fourier transform to extract the frequency information of the time series data. Based on the period length of the data, identify the main periodic features by analyzing the frequency component with the largest amplitude, and this periodic feature is the preprocessed text file.
[0125] In the virtual user behavior prediction analysis algorithm of this application, the calculation metrics of the user include main calculation metrics and auxiliary calculation metrics. The main calculation metrics include the file similarity processed by the analyzed user, the event similarity processed by the analyzed user, and the periodic similarity processed by the analyzed user. The auxiliary calculation metrics include whether the analyzed user has completed the event, the amount of events currently completed by the analyzed user, and the ratio of the completed amount to the accessed amount of the analyzed user.
[0126] 1. The calculation rules of the main metrics are as follows:
[0127] Extract the information such as the title and keywords in the current file, perform a similarity comparison analysis with the keywords of historical files in the database, take the maximum similarity value of each group of keywords in the current file and the keywords of historical files, and then take the average of multiple maximum values as the similarity value of the file. Files with a similarity value of 0 do not participate in the calculation. Bind the historical files to each node information to obtain the events corresponding to the files related to or similar to the current file, exclude the data of the current node, and extract the data of file similarity and average value for each node.
[0128] The similarity calculation formula for each file:
[0129] cosθ = (A·B) / (||A|| × ||B||)
[0130] cosθ is the similarity value; A·B represents the dot product of vectors A and B, and ||A|| and ||B|| are the norms of vectors A and B (i.e., the lengths of the vectors), which can be obtained by taking the square root of the sum of the squares of each element. The value range of cosine similarity is between -1 and 1. When the directions of two vectors are exactly the same (the included angle is 0°), the cosine similarity is 1; when the directions of two vectors are exactly opposite (the included angle is 180°), the cosine similarity is -1; when two vectors are orthogonal (the included angle is 90° or 270°), the cosine similarity is 0. Therefore, the closer the value of cosine similarity is to 1, the more consistent the directions of the two vectors are, and the more similar they are; on the contrary, the closer the value is to -1, the more opposite the directions of the two vectors are, and the less similar they are. Events with the same similarity belong to the same category.
[0131] Retrieve the approval processing information of the tasks sent by the user at the current node, match the similarity according to the key information and the processing information in the historical data, exclude the data of the current user, take the maximum similarity value of each group of keywords in the current approval and the keywords of historical files, and then take the average of multiple maximum values as the similarity value of the approval. Files with a similarity value of 0 do not participate in the calculation. Bind the historical data to each node information and take the average value of each node.
[0132] Query the historical cycle information record of the user at the current node and exclude the current user selection from the calculation.
[0133] 2. Auxiliary index calculation rules:
[0134] Obtain whether the current user is online. If online, add 1 to the score;
[0135] Obtain the events that the current user has visited. If there are multiple events, then add 1 to each event. Events with the same similarity belong to the same category;
[0136] Query the number of events completed by the user currently. If there are multiple completed events, then retrieve all completed events. Events with the same similarity belong to the same category;
[0137] Completion access ratio (T) = Completed events (ST) / Sum of accessed events (AT)
[0138] Filter such that the completion access ratio (T) is 1, the completed events (ST) is 1, and the sum of accessed events (AT) is 1;
[0139] If the completion access ratio (T) is higher, it indicates that the user is more real and the prediction is more accurate;
[0140] Detailed calculation method:
[0141] Events with the same similarity belong to one category
[0142]
[0143] Calculate the value at A. The value at A depends on ra and pa. ra represents similarity, and pa represents the completion access ratio;
[0144] Prediction algorithm calculation formula:
[0145]
[0146] x represents a completed event, and y represents the events viewed per day;
[0147] First, calculate the average value of X, the average value of Y, the score (completion access ratio) of the X event occurring in the most recent n days, x / n is the average value of X; the number of times the Y event is viewed in the most recent n days, y / n is the average value of Y;
[0148] Calculate the regression coefficient f(x,y): Use the above formula to calculate the regression coefficient f(x,y), and this coefficient represents the strength of the linear relationship between X and Y;
[0149] Finally, using the obtained regression coefficient f(x,y), as well as the previously calculated average values, a linear regression equation can be constructed y= b x+ay= b x+a , where
[0150] Based on the above calculations, for each virtual identity, in the n-day cycle, prediction information such as public opinion expression, real-time comments, and contraband items is obtained, and a top-N sorted result set is generated. For the subsequent behavior of the virtual identity, predictions and judgments are made based on the top results in the result set. Based on the prediction and judgment results of each virtual identity, geographical locations are supplemented to achieve the purpose of this prediction and judgment calculation.
[0151] The network virtual user behavior analysis system and its analysis method of the present invention have important application values, can provide strong support for security monitoring, etc., and can be applied to aspects such as security guarantee of major events and related places, analysis of dangerous behaviors of network virtual personnel, and periodic attention to dangerous behaviors.
[0152] Based on the open-source data stream processing algorithm, this algorithm model supports Flink real-time computing and Spark real-time computing.
[0153] It should be emphasized that the term "including / comprising" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0154] In addition, the method of the present invention is not limited to being executed in the time sequence described in the specification, and can also be executed in other time sequences, in parallel or independently. Therefore, the execution sequence of the method described in this specification does not limit the technical scope of the present invention.
[0155] Although the present invention has been disclosed above through the description of specific embodiments of the present invention, it should be understood that all the above embodiments and examples are exemplary, not restrictive. Those skilled in the art can design various modifications, improvements or equivalents to the present invention within the spirit and scope of the appended claims. These modifications, improvements or equivalents should also be considered to be included within the protection scope of the present invention.
Claims
1. A network virtual identity behavior prediction and analysis method, characterized in that: include: Calculate the similarity of files processed by the analyzed user, events processed by the analyzed user, and periods processed by the analyzed user based on the BERT model; The similarity of files processed by the analyzed user, the similarity of events processed by the analyzed user, and the similarity of periods processed by the analyzed user are classified and aggregated according to the virtual identity; Perform regression prediction calculations on the classified and aggregated virtual identity data to predict the behavioral trends and potential risks of each virtual identity.
2. The network virtual identity behavior prediction and analysis method according to claim 1 is characterized by: Also includes: According to the results of regression prediction, select virtual identities with higher risks to obtain the topN result set; The location information is supplemented for the virtual identities in the topN result sets, and finally the virtual identity personnel with dangerous behaviors are obtained.
3. The network virtual identity behavior prediction and analysis method according to claim 1 is characterized by: The calculation of the similarity of files processed by the analyzed user based on the BERT model specifically includes: Step 1, data collection and preprocessing: collect a set of text files related to user behavior, extract titles and keywords from each text file in the text file set as preprocessed text; Step 2: read one of the text files in the text file set as the current text file, and compare and analyze the similarity between the title and keywords of the current text file and the keywords of the historical files in the database through the BERT model; Step 3, for each group of keywords in the current text file, after comparing with the keywords in the historical files, select the maximum similarity; Step 4, calculate the file similarity value: take the average of multiple maximum values as the similarity value of the current text file; but if the similarity value is 0, it will not be included in the calculation; Step 5: Bind the historical files to each user. At the same time, exclude the current user data during the analysis process and extract the data of similarity and average value of each user file. Step 6, repeating steps 2 to 5, to obtain the similarity and average value data of all text files in the text file set as the file similarity.
4. The network virtual identity behavior prediction and analysis method according to claim 1 is characterized by: The similarity of events processed by the analyzed user is calculated based on the BERT model, including: Step 11, describing the event in text form, extracting text content related to the event, and then removing noise from the text content to obtain a cleaned text; extracting valuable feature words from the cleaned text to form a preprocessed text file; valuable feature words are feature words that measure their importance in the text content, including their frequency of appearance in a single text file greater than a preset value, and their rarity in the entire text collection within a preset threshold range; Step 12, read one of the text files in the preprocessed text file set as the current text file, and compare and analyze the similarity between the feature words of the current text file and the feature words of the historical files in the database through the BERT model; during the comparison process, the similarity between the feature words in the current text file and all the feature words in the historical files will be calculated one by one; Step 13, for each group of feature words in the current text file, after comparing with the feature words in the historical file, select the maximum value of the similarity; the maximum value represents the closeness between the group of feature words and the most similar feature words in the historical file. In this way, the strongest association between the feature words in the current file and the feature words in the historical file can be found; Step 14, calculate the file similarity value: take the average of multiple maximum values as the similarity value of the current text file; but if the similarity value is 0, it will not be included in the calculation; Step 15, bind the historical files to each user, exclude the current user data in the analysis process, and extract the data of the similarity value and average value of each user file; Step 16, repeating steps 12 to 15, obtaining the similarity values and average values of all text files in the text file set as event similarity.
5. The network virtual identity behavior prediction and analysis method according to claim 1 is characterized by: The calculation of the similarity of the cycles processed by the analyzed user according to the BERT model specifically includes: Step 21, time series data extraction: for events with periodicity, extract time series data, and use interpolation method to fill in missing data to ensure the accuracy and completeness of time series data; periodic feature extraction: determine the period length of the data, and extract periodic features based on the period length of the data; the periodic features form a preprocessed text file; Step 22, read one of the text files in the preprocessed text file set as the current text file, and compare and analyze the similarity between the periodic features of the current text file and the features of the historical files in the database through the BERT model; during the comparison process, the similarity between the periodic features in the current text file and all the features in the historical files will be calculated one by one; Step 23, for each set of periodic features of the current text file, after comparing with the historical file features, select the maximum value of similarity; the maximum value represents the closeness between the set of periodic features and the most similar features in the historical file. In this way, the strongest correlation between the current file features and the historical file features can be found; Step 24, calculate the file similarity value: take the average of multiple maximum values as the similarity value of the current text file; but if the similarity value is 0, it will not be included in the calculation; Step 25, bind the historical files to each user, exclude the current user data in the analysis process, and extract the data of the similarity value and average value of each user file; Step 26, repeating steps 22 to 25, obtaining the similarity values and average values of all text files in the text file set as periodic similarity.
6. The network virtual identity behavior prediction and analysis method according to claim 3 is characterized by: In step 4, the similarity calculation formula for each user file is: cosθ=(A·B) / (||A||×||B||); Where cosθ is the cosine similarity value, θ represents the angle between vector A and vector B; A·B represents the dot product of vector A and vector B, and ||A|| and ||B|| are the norms of vector A and vector B, respectively.
7. The network virtual identity behavior prediction and analysis method according to claim 1 is characterized by: In addition to the above-mentioned similarity of files processed by the analyzed user, similarity of events processed by the analyzed user, and similarity of periods processed by the analyzed user as the main calculation indicators, auxiliary calculation indicators include whether the analyzed user has completed the event, the number of events currently completed by the analyzed user, and the ratio of the completed number to the access number of the analyzed user.
8. The network virtual identity behavior prediction and analysis method according to claim 1 is characterized by: The calculation rules of the auxiliary calculation index are as follows: Get whether the current user is online. If online, the score increases by 1. Get the events that the current user has visited. If there are multiple events, add 1 to each event, and the events with the same similarity are classified into the same category; Query the number of events that the user has completed. If there are multiple completed events, take out all completed events and classify them as events with the same similarity. Completion access ratio (T) = completion events (ST) / total number of visited events (AT); The filtering completed access ratio (T) is 1, the completed event (ST) is 1, and the visited event sum (AT) is 1; If the completed visit ratio (T) is higher, it means that the user is real and the prediction is more accurate.
9. The network virtual identity behavior prediction and analysis method according to claim 1 is characterized by: The prediction algorithm calculation formula of this prediction analysis method is: x represents a completed event, y represents the events viewed every day; X is the set of events x completed every day, Y is the set of events viewed every day; n is the number of days, x n Indicates the event completed on day n, y n Indicates the event viewed on day n; First calculate the mean of X and the mean of Y: x / n is the mean of X, y / n is the mean of Y; Then use the above formula to calculate the regression coefficient f(x,y); Finally, the linear regression equation is constructed using the obtained regression coefficient f(x,y) and the mean value of X and the mean value of Y. y= b x+a ,in Based on the above calculations, we can obtain the predicted information of public opinion, real-time comments, prohibited items, etc. for each virtual identity in an n-day period, and generate a topN sorted result set.
10. A network virtual identity behavior prediction and analysis system, characterized in that: include: A similarity calculation module is used to calculate the similarity of files processed by the analyzed user, the similarity of events processed by the analyzed user, and the similarity of periods processed by the analyzed user according to the BERT model; A classification and aggregation module is used to classify and aggregate the similarities of files processed by the analyzed user, the similarities of events processed by the analyzed user, and the similarities of periods processed by the analyzed user according to the virtual identity; A regression prediction calculation module is used to perform regression prediction calculation on the classified and aggregated virtual identity data to predict the behavior trend and potential risks of each virtual identity; The network virtual identity behavior prediction and analysis system is used to execute the network virtual identity behavior prediction and analysis method described in any one of claims 1-9.
Citation Information
Patent Citations
Behavior trajectory prediction method and device, electronic equipment and storage medium
CN115345390A