Short video traffic anomaly analysis system based on big data analysis
Through a short video traffic anomaly analysis system based on big data analysis, combined with a variety of data analysis methods, the problem of difficulty in identifying complex cheating behaviors in the existing technology is solved, and higher abnormal detection accuracy and data quality are achieved.
Patent Information
- Application Number
- CN202510251584.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-24
AI Technical Summary
The existing short video traffic anomaly detection system relies on methods based on rules and statistical analysis, and it is difficult to effectively identify complex and highly concealed cheating methods, and it is easy to miss reports.
A short video traffic anomaly analysis system based on big data analysis is adopted, and a variety of data analysis methods such as data collection and preliminary screening, user behavior analysis, comment content analysis, and social relationship analysis are used to comprehensively analyze user behavior and identify abnormal users.
It improves the accuracy of abnormal detection, can effectively identify complex traffic abnormalities and cheating behaviors, reduces the risk of missed and false alarms, and improves data quality and effectiveness of abnormal detection.
Smart Images

Figure CN120201214A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and particularly to a short video traffic anomaly analysis system based on big data analysis. Background Technique
[0002] With the rapid development of the short video industry, short video platforms have attracted the participation of a large number of users and creators, and the massive data generated on the platforms is growing increasingly large. The speed of video content dissemination and the frequency of user interactions have made short video platforms a highly dynamic and complex environment. In such an environment, traffic anomaly problems often lead to a decline in the user experience of the platform and may even affect the reputation and profit model of the platform. Therefore, how to efficiently and accurately analyze short video traffic anomalies, identify malicious traffic brushing, cheating behaviors, and other improper means has become an important challenge faced by short video platforms.
[0003] Currently, existing short video traffic anomaly detection systems rely on rule-based and statistical analysis methods. These systems monitor traffic data on the platform in real time and identify anomalies by setting certain rules or thresholds. Although this method can identify traffic anomalies to a certain extent, it is unable to cope when faced with increasingly complex cheating means. For example, behaviors such as malicious traffic brushing, forged likes, and comments are not only highly concealed but may also be disguised as normal user behaviors through means such as machine learning and automated scripts. Therefore, traditional rule-based and threshold-based detection methods cannot effectively identify these complex and highly concealed cheating means and are prone to false negatives. Summary of the Invention
[0004] The purpose of the present invention is to provide a short video traffic anomaly analysis system based on big data analysis to solve the problems raised in the above background technique.
[0005] To solve the above technical problems, the present invention provides the following technical solutions:
[0006] A short video traffic anomaly analysis system based on big data analysis includes: a data collection and preliminary screening module, a user behavior analysis module, a comment content analysis module, a social relationship analysis module, and an abnormal user identification and review module;
[0007] The data collection and preliminary screening module obtains the traffic data and user behavior data of the short video to be analyzed within a selected time period, and according to preset rules or thresholds, screens out the traffic data to be analyzed and extracts the target user numbers and the corresponding target user behavior data;
[0008] The user behavior analysis module analyzes the behavior patterns of target users within the selected time period based on the target user behavior data, and conducts a singularity evaluation of the behavior patterns of target users according to the analysis results of the behavior patterns;
[0009] The comment content analysis module extracts the comment information of the target user in the short video to be analyzed, performs keyword extraction and semantic analysis on the comment information and the short video to be analyzed respectively, and evaluates the content compliance between the comment information of the target user and the content of the short video to be analyzed;
[0010] The social relationship analysis module constructs a target user set according to the target user numbers corresponding to the target user behavior data; for each target user, obtains the corresponding social relationship data, thereby constructing a social relationship graph of the target user; based on the social relationship graph, analyzes the social complexity of the target user of the short video to be analyzed;
[0011] The abnormal user identification and review module identifies abnormal users based on the analysis results of the behavioral pattern singularity, content compliance, and correlation degree of the target user; outputs the abnormal user numbers to relevant personnel, who conduct reviews and screen out abnormal traffic data according to the review results.
[0012] Furthermore, the data collection and preliminary screening module includes a data collection unit and a data preliminary screening unit;
[0013] The data collection unit is used to collect the traffic data and user behavior data of the short video to be analyzed within a selected time period. The traffic data refers to the statistical information of various interaction and viewing behaviors of the short video to be analyzed; the user behavior data refers to the record of various behavior information of the user during the process of watching the short video to be analyzed;
[0014] The data preliminary screening unit preliminarily screens out the traffic data to be analyzed of the short video to be analyzed within a selected time period according to preset rules or thresholds, and based on the traffic data to be analyzed, extracts the corresponding target user numbers and target user behavior data; the preset rules or thresholds refer to the common methods in the prior art for identifying abnormal traffic data. For example, setting a certain traffic growth rate threshold, such as marking as abnormal when the traffic increases by more than 50% or 100% compared to the normal traffic; comparing the number of views, likes, or comments of the user within a specific time period with their historical behavior data to detect behaviors that deviate from the normal range; monitoring the click-through rate (CTR) and conversion rate (CVR) of the video, and marking as abnormal traffic if these indicators are significantly higher than the historical data average or industry benchmark value. The traffic data to be analyzed refers to the traffic data that is preliminarily judged to be normal under the preset rules or thresholds. These data represent that according to the preset rules or thresholds, it is preliminarily judged that the short video to be analyzed has normal traffic and real user behavior data within a selected time period, but has not been filtered for malicious brushing, cheating behaviors, and other improper means. Therefore, these traffic data may contain some abnormal behaviors or improper traffic that require further analysis and processing.
[0015] Further, the user behavior analysis module includes a standard user data analysis unit, a behavior pattern recognition unit, and a behavior pattern singularity evaluation unit;
[0016] The standard user data analysis unit extracts standard user behavior data from the database and preprocesses the standard user behavior data; analyzes the preprocessed standard user behavior data to obtain a list of behavior patterns of standard users; the behavior pattern recognition unit preprocesses the target user behavior data of the target user within a selected time period, and analyzes the preprocessed target user behavior data with the standard user behavior data in the list of behavior patterns of standard users to identify the behavior pattern of the target user; the behavior pattern singularity evaluation unit calculates the behavior pattern singularity score of the target user within the selected time period according to the analysis result of the behavior pattern of the target user.
[0017] Further, the standard user data analysis unit obtains a list of behavior patterns of standard users, and the specific analysis process is as follows:
[0018] Obtain the standard user behavior data of a selected period from the database. The standard user refers to a user who has been checked by relevant personnel and has no abnormal behavior; extract the time series data from the standard user behavior data and perform standardization processing to obtain a behavior data sequence B(t), and B(t) = {b1(t), b2(t),..., bn(t)}, where B(t) represents the behavior data sequence of the standard user at time t, b1(t) represents the intensity or frequency of the first behavior, b2(t) represents the intensity or frequency of the second behavior, and so on, bn(t) represents the intensity or frequency of the nth behavior; n represents the number of types of all behaviors of the standard user; calculate the similarity between the behavior data sequences at different times, and use a clustering algorithm to divide the behavior data sequences at different times into different behavior patterns according to the similarity between the behavior data sequences at different times, so as to obtain a list of behavior patterns BD of standard users, and each behavior pattern p corresponds to a number of behavior data sequences of standard users.
[0019] Further, the specific content of the behavior pattern recognition unit for identifying the behavior pattern of the target user within the selected time period includes:
[0020] Obtain the target user behavior data of the target user within the selected time period, convert the target user behavior data into time series data, and perform standardization processing to obtain a behavior data sequence X(t), and X(t) = {x1(t), x2(t),..., xn(t)}, where X(t) represents the behavior data sequence of the target user at time t, xi(t) represents the intensity or frequency of the i-th behavior, such as viewing duration, number of likes, etc.; t takes values from 1 to n, and n represents the number of types of behaviors;
[0021] The behavior data sequence X(t) of the target user at different times is successively calculated for similarity with the behavior data sequence B(t) of each behavior pattern p in the behavior pattern list BD of the standard user, and the average value is obtained, so as to obtain the similarity Sp between the behavior data sequence X(t) of the target user at different times and the behavior data sequence B(t) corresponding to each behavior pattern p in the behavior pattern list BD of the standard user; the behavior pattern p with the largest similarity Sp is selected as the behavior pattern of the behavior data sequence X(t) of the target user, and all the behavior data sequences X(t) of the target user in the selected time period are traversed, so as to obtain the behavior pattern of the target user in the selected time period.
[0022] Further, the behavior pattern singularity evaluation unit calculates the behavior pattern singularity score of the target user in the selected time period, and the specific calculation process is as follows:
[0023] According to the time sequence, the behavior patterns of the behavior data sequence X(t) at different times are summarized, the number of types of the behavior patterns of the target user in the selected time period is counted, and the number of types of the behavior patterns of the target user in the selected time period is used as the behavior consistency measurement parameter M; the deviations of the behavior patterns of two adjacent behavior data sequences X(t) are compared in turn, and all the behavior data sequences X(t) of the target user in the selected time period are traversed to calculate the average deviation, so as to obtain the behavior pattern change measurement parameter D of the target user in the selected time period;
[0024] According to the behavior consistency measurement parameter M and the change measurement parameter D of the target user in the selected time period, the behavior pattern singularity score C of the target user in the selected time period is calculated, and C = k / (α×M + β×D), where k represents the adjustment parameter, α and β represent the weight coefficients, and α + β = 1.
[0025] Further, the comment content analysis module includes a text preprocessing and keyword analysis unit and a semantic analysis and compliance evaluation unit;
[0026] The text preprocessing and keyword analysis unit extracts the comment information of the target user under the short video to be analyzed, and performs text preprocessing on the comment information of the target user under the short video to be analyzed, so as to obtain the text set W of the comment information of the target user; perform format conversion on the short video to be analyzed, and convert the video data into text data, so as to obtain the corresponding text set V; wherein the process of converting video data into text data is as follows: extract voice information through speech recognition, then combine image analysis technology to obtain visual content, and finally use natural language generation technology to convert these information into text. For the text set W of the comment information of the target user, count the proportion R of repeated comments and the proportion N of meaningless characters; extract the keywords with the highest TF-IDF weights from the text set W and the text set V respectively, and denote them as the keyword set Kw and the keyword set Kv; calculate the keyword matching degree G between the keyword set Kw and the keyword set Kv, and the specific calculation formula is: G = |Kw ∩ Kv| / |Kw ∪ Kv|;
[0027] The semantic analysis and compliance evaluation unit maps the content of the comment information of the target user and the content of the short video to be analyzed into semantic vectors respectively through a pre-trained NLP model, so as to be expressed as: Ew and Ev; according to the semantic vectors Ew and Ev, calculate the semantic correlation degree Y between the two, and the specific calculation formula is: Y = (Ew · Ev) / (||Ew|| ||Ev||); combine the proportion R of repeated comments, the proportion N of meaningless characters, the keyword matching degree G and the semantic correlation degree Y, and comprehensively evaluate the compliance degree F between the comment information of the target user and the content of the short video to be analyzed, and F = w1 × G + w2 × Y - w3 × (R + N), where w1, w2 and w3 represent weight coefficients, and w1 + w2 + w3 = 1.
[0028] Furthermore, the social relationship analysis module includes a social relationship graph construction unit and a social complexity analysis unit;
[0029] The social relationship graph construction unit constructs a target user set U according to the target user numbers corresponding to the target user behavior data, and U = {u1, u2,..., um}, where uj represents the jth target user number, j takes values from 1 to m, and m represents the number of target users; for each target user, obtain the corresponding social relationship data and construct a social relationship graph S(V, E), where V represents nodes, that is, the relevant users XU extracted from the social relationship data corresponding to the target user, which can include friend lists, followed users, comment interactions, like interactions, etc.; E represents edges, if there is social relationship data between two users, an edge is established in the social relationship graph, and the edge only represents the existence of a social relationship, and the length of the edge has no meaning;
[0030] The social complexity analysis unit extracts all social paths from the social relationship graph S(V, E) of each target user, so as to obtain the social path set Pa of the corresponding target user; the social path represents the path starting from the target user to any other relevant user; for each target user of the short video to be analyzed, combined with the corresponding social path set Pa, the social complexity Z of the target user is analyzed. The specific calculation formula is: Z = [1 / N(Pa)]·Σpk∈PaL(pk), where N(Pa) represents the number of social paths in the social path set Pa, pk represents the k-th social path, and L(pk) represents the number of users in the k-th social path.
[0031] Further, the abnormal user identification and review module includes an abnormal user identification unit and a review unit;
[0032] The abnormal user identification unit comprehensively analyzes based on the behavioral pattern singularity, content compliance, and social relationship complexity of the target user, so as to obtain the comprehensive abnormal user score H, and H = a×C + b×(1 - F) + c×(1 - Z); compare the comprehensive abnormal user score H of the target user of the short video to be analyzed with the threshold H0. If H > H0, the target user is marked as an abnormal user; where the threshold H0 is obtained by performing the same analysis process as the target user based on the standard users in the database, and the minimum value of the comprehensive abnormal user score corresponding to the standard users is selected as the threshold H0;
[0033] The review unit obtains the corresponding abnormal user number according to the identification result of the abnormal user, outputs the abnormal user number to the relevant personnel, and the relevant personnel conduct a review, and screen out the abnormal traffic data of the short video to be analyzed according to the review result.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows: Compared with traditional rule- and statistical analysis-based methods, this system adopts a variety of data analysis methods (user behavior analysis, comment content analysis, social relationship analysis, etc.), which can comprehensively analyze user behavior from different perspectives and improve the accuracy of anomaly detection; this multi-dimensional analysis method enables the system to more effectively identify complex traffic anomalies and cheating behaviors. By dynamically identifying and evaluating the singularity of the behavior patterns of target users, abnormal users whose behaviors are significantly different from those of standard users can be identified in a timely manner; this behavior pattern-based analysis method enables the system to more flexibly respond to different cheating techniques and reduce the risks of missed reports and false alarms. The comment content analysis module conducts in-depth text and semantic analysis of user comments to ensure accurate evaluation of the compliance between the comment content and the short video content; this deep semantic understanding ability enables the system to identify forged or irrelevant comments, thereby further improving data quality and the accuracy of anomaly detection. By establishing a social relationship graph of users through the social relationship analysis module and analyzing social complexity, potential cheating behaviors are searched from the perspective of the social network, and abnormal features that are difficult to identify by rules and thresholds can be captured; this unique analysis dimension helps to comprehensively improve the effectiveness of anomaly detection. The abnormal user identification and review module introduces a comprehensive scoring mechanism. By combining the singularity of behavior patterns, content compliance, and social relationship complexity, a comprehensive scoring index is formed; this mechanism can achieve a comprehensive evaluation of target users, making the identification of abnormal users more rigorous. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings:
[0036] Figure 1 is a schematic diagram of the modules of a short video traffic anomaly analysis system based on big data analysis according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] Please refer to Figure 1 , the present invention provides the following technical solutions:
[0039] A short video traffic anomaly analysis system based on big data analysis, comprising: a data collection and preliminary screening module, a user behavior analysis module, a comment content analysis module, a social relationship analysis module, and an abnormal user identification and review module;
[0040] The data collection and preliminary screening module obtains the traffic data and user behavior data of the short video to be analyzed within a selected time period, and according to preset rules or thresholds, screens out the traffic data to be analyzed and extracts the target user numbers and the corresponding target user behavior data;
[0041] The user behavior analysis module analyzes the behavior patterns of target users within a selected time period based on the target user behavior data, and conducts a singularity evaluation of the behavior patterns of target users according to the analysis results of the behavior patterns;
[0042] The comment content analysis module extracts the comment information of target users on the short video to be analyzed, conducts keyword extraction and semantic analysis on the comment information and the short video to be analyzed respectively, and evaluates the content compliance between the comment information of target users and the content of the short video to be analyzed;
[0043] The social relationship analysis module constructs a target user set according to the target user numbers corresponding to the target user behavior data; for each target user, obtains the corresponding social relationship data, thereby constructing a social relationship graph of the target user; based on the social relationship graph, analyzes the social complexity of the target users of the short video to be analyzed;
[0044] The abnormal user identification and review module identifies abnormal users based on the analysis results of the singularity, content compliance and correlation degree of the behavior patterns of target users; outputs the abnormal user numbers to relevant personnel, who conduct reviews, and screens out abnormal traffic data according to the review results.
[0045] The data collection and preliminary screening module includes a data collection unit and a data preliminary screening unit;
[0046] The data collection unit is used to collect the traffic data and user behavior data of the short video to be analyzed within a selected time period. The traffic data refers to the statistical information of various interaction and viewing behaviors of the short video to be analyzed; the user behavior data refers to the records of various behavior information of users during the process of watching the short video to be analyzed;
[0047] The preliminary data screening unit preliminarily screens out the traffic data to be analyzed of the short video to be analyzed within the selected time period according to preset rules or thresholds, and based on the traffic data to be analyzed, extracts the corresponding target user numbers and target user behavior data; the preset rules or thresholds refer to common methods in the prior art for identifying abnormal traffic data, such as setting a certain traffic growth rate threshold, for example, marking as abnormal when the traffic increases by more than 50% or 100% compared with the normal traffic; by comparing the number of views, likes or comments of users within a specific time period with their historical behavior data, detecting behaviors that deviate from the normal range; monitoring the click-through rate (CTR) and conversion rate (CVR) of the video, and if these indicators are significantly higher than the historical data average or industry benchmark value, marking them as abnormal traffic. The traffic data to be analyzed refers to the traffic data that is preliminarily judged to be normal under the preset rules or thresholds. These data represent that, according to the preset rules or thresholds, it is preliminarily judged that the short video to be analyzed has normal traffic and real user behavior data within the selected time period, but it has not been filtered for malicious traffic brushing, cheating behaviors and other improper means. Therefore, these traffic data may contain some abnormal behaviors or improper traffic that need further analysis and processing.
[0048] The user behavior analysis module includes a standard user data analysis unit, a behavior pattern recognition unit, and a behavior pattern singularity evaluation unit;
[0049] The standard user data analysis unit extracts standard user behavior data from the database and preprocesses the standard user behavior data; analyzes the preprocessed standard user behavior data to obtain a list of standard user behavior patterns; the behavior pattern recognition unit preprocesses the target user behavior data of the target user within the selected time period, and analyzes the preprocessed target user behavior data with the standard user behavior data in the list of standard user behavior patterns to identify the behavior pattern of the target user; the behavior pattern singularity evaluation unit calculates the behavior pattern singularity score of the target user within the selected time period according to the analysis result of the behavior pattern of the target user.
[0050] The standard user data analysis unit obtains a list of standard user behavior patterns. The specific analysis process is as follows:
[0051] Obtain the standard user behavior data for the selected period from the database. The standard user refers to a user who has been checked by relevant personnel and has no abnormal behavior. Extract the time series data from the standard user behavior data and perform standardization processing to obtain the behavior data sequence B(t), and B(t) = {b1(t), b2(t),..., bn(t)}, where B(t) represents the behavior data sequence of the standard user at time t, b1(t) represents the intensity or frequency of the first type of behavior, b2(t) represents the intensity or frequency of the second type of behavior, and so on. bn(t) represents the intensity or frequency of the nth type of behavior; n represents the number of types of all behaviors of the standard user; Calculate the similarity between the behavior data sequences at different times, and according to the similarity between the behavior data sequences at different times, use the clustering algorithm to divide the behavior data sequences at different times into different behavior patterns, so as to obtain the behavior pattern list BD of the standard user, and each behavior pattern p corresponds to the behavior data sequences of several standard users.
[0052] In this embodiment, to calculate the similarity between the behavior data sequences at different times, common similarity calculation methods are used, such as cosine similarity or Euclidean distance. By calculating the similarity between the behavior data sequences at different time points, it can be understood which time points have similar behavior patterns, which is helpful for subsequent clustering analysis.
[0053] According to the similarity between the behavior data sequences at different times, use the clustering algorithm to divide the behavior data sequences into different behavior patterns. Common clustering algorithms include K-means clustering, DBSCAN, or hierarchical clustering.
[0054] Suppose the K-means clustering algorithm is used to divide the behavior patterns. The main steps of the K-means clustering algorithm are as follows:
[0055] 1. Initialize the centers: Randomly select K initial clustering centers (i.e., the center points of the behavior patterns);
[0056] 2. Assign clusters: Assign each behavior data sequence to the nearest clustering center;
[0057] 3. Update the clustering centers: Update the clustering centers according to the mean of the behavior data sequences in each cluster;
[0058] 4. Repeat steps 2 and 3: Until the clustering result converges, that is, the clustering centers no longer change.
[0059] The specific content of the behavior pattern recognition unit for recognizing the behavior pattern of the target user within the selected time period includes:
[0060] Obtain the target user behavior data of the target user within the selected time period, convert the target user behavior data into time series data, and perform normalization processing to obtain the behavior data sequence X(t), and X(t) = {x1(t), x2(t),..., xn(t)}, where X(t) represents the behavior data sequence of the target user at time t, and xi(t) represents the intensity or frequency of the i-th behavior, such as viewing duration, number of likes, etc.; t takes values from 1 to n, and n represents the number of types of behaviors;
[0061] Calculate the similarity between the behavior data sequences X(t) of the target user at different times and the behavior data sequences B(t) of each behavior pattern p in the behavior pattern list BD of the standard user in turn, and calculate the average value, so as to obtain the similarity Sp between the behavior data sequence X(t) of the target user at different times and the behavior data sequence B(t) corresponding to each behavior pattern p in the behavior pattern list BD of the standard user; Select the behavior pattern p with the largest similarity Sp as the behavior pattern of the behavior data sequence X(t) of the target user, and traverse all the behavior data sequences X(t) of the target user within the selected time period, so as to obtain the behavior pattern of the target user within the selected time period.
[0062] The behavior pattern singularity evaluation unit calculates the behavior pattern singularity score of the target user within the selected time period, and the specific calculation process is as follows:
[0063] Summarize the behavior patterns of the behavior data sequences X(t) at different times in chronological order, count the number of types of behavior patterns of the target user within the selected time period, and use the number of types of behavior patterns of the target user within the selected time period as the behavior consistency measurement parameter M; Compare the deviations of the behavior patterns of two adjacent behavior data sequences X(t) in turn, traverse all the behavior data sequences X(t) of the target user within the selected time period, and calculate the average deviation, so as to obtain the change measurement parameter D of the behavior pattern of the target user within the selected time period;
[0064] According to the behavior consistency measurement parameter M and the change measurement parameter D of the target user within the selected time period, calculate the behavior pattern singularity score C of the target user within the selected time period, and C = k / (α×M + β×D), where k represents the adjustment parameter, α and β represent the weight coefficients, and α + β = 1.
[0065] In this embodiment, assume that there is a target user's behavior data sequence X(t) within the selected time period, and the data is as shown below, including the behavior data at 5 time points:
[0066] At time point t1, the behavior data is: the frequency of behavior 1 is 3, the frequency of behavior 2 is 5, and the frequency of behavior 3 is 2;
[0067] At time point t2, the behavioral data is as follows: the frequency of behavior 1 is 4, the frequency of behavior 2 is 6, and the frequency of behavior 3 is 3;
[0068] At time point t3, the behavioral data is as follows: the frequency of behavior 1 is 2, the frequency of behavior 2 is 4, and the frequency of behavior 3 is 1;
[0069] At time point t4, the behavioral data is as follows: the frequency of behavior 1 is 3, the frequency of behavior 2 is 5, and the frequency of behavior 3 is 2;
[0070] At time point t5, the behavioral data is as follows: the frequency of behavior 1 is 5, the frequency of behavior 2 is 7, and the frequency of behavior 3 is 4;
[0071] According to the behavior pattern recognition unit, the following results are obtained:
[0072] At time point t1, the behavior pattern is pattern 1; at time point t2, the behavior pattern is pattern 1; at time point t3, the behavior pattern is pattern 2; at time point t4, the behavior pattern is pattern 1; at time point t5, the behavior pattern is pattern 3;
[0073] Count the number of types of behavior patterns of the target user within the selected time period. The behavior patterns of the target user within the time period are pattern 1, pattern 2, and pattern 3, with a total of 3 different behavior patterns; therefore, the value of the behavior consistency measurement parameter M is 3.
[0074] Calculate the change measurement parameter D of the behavior pattern of the target user within the selected time period. It is necessary to calculate the deviation of the behavior pattern between two adjacent time points; assume that the change of the behavior pattern is represented as the difference between patterns. If the behavior patterns are the same, the deviation is 0; if the behavior patterns are different, the deviation is 1;
[0075] From time point t1 to t2, the behavior pattern does not change, and the deviation is 0;
[0076] From time point t2 to t3, the behavior pattern changes from pattern 1 to pattern 2, and the deviation is 1;
[0077] From time point t3 to t4, the behavior pattern changes from pattern 2 to pattern 1, and the deviation is 1;
[0078] From time point t4 to t5, the behavior pattern changes from pattern 1 to pattern 3, and the deviation is 1;
[0079] Therefore, the deviation sequence is: 0, 1, 1, 1. To calculate the change measurement parameter D, calculate the average value of these deviations:
[0080] D = (0 + 1 + 1 + 1) / 4 = 3 / 4 = 0.75;
[0081] Calculate the behavioral pattern singularity score C of the target user within the selected time period. Assume that k = 10, α = 0.7, and β = 0.3. Then, the behavioral pattern singularity score C is calculated as: C = 10 / (0.7×3 + 0.3×0.75) = 4.3.
[0082] The comment content analysis module includes a text preprocessing and keyword analysis unit and a semantic analysis and compliance evaluation unit;
[0083] The text preprocessing and keyword analysis unit extracts the comment information of the target user under the short video to be analyzed, performs text preprocessing on the comment information of the target user under the short video to be analyzed, so as to obtain the text set W of the comment information of the target user; performs format conversion on the short video to be analyzed, converts the video data into text data, so as to obtain the corresponding text set V, where the process of converting video data into text data is: extracts voice information through speech recognition, then combines image analysis technology to obtain visual content, and finally uses natural language generation technology to convert this information into text. For the text set W of the comment information of the target user, statistically calculate the proportion R of repeated comments and the proportion N of meaningless characters; extract the keywords with the highest TF-IDF weights from the text set W and the text set V respectively, and denote them as the keyword set Kw and the keyword set Kv; calculate the keyword matching degree G between the keyword set Kw and the keyword set Kv, and the specific calculation formula is: G = |Kw ∩ Kv| / |Kw ∪ Kv|;
[0084] The semantic analysis and compliance evaluation unit maps the content of the comment information of the target user and the content of the short video to be analyzed into semantic vectors respectively through a pre-trained NLP model, so as to be represented as: Ew and Ev; according to the semantic vectors Ew and Ev, calculate the semantic correlation degree Y between the two, and the specific calculation formula is: Y = (Ew · Ev) / (||Ew||||Ev||); combining the proportion R of repeated comments, the proportion N of meaningless characters, the keyword matching degree G, and the semantic correlation degree Y, comprehensively evaluate the compliance degree F between the comment information of the target user and the content of the short video to be analyzed, and F = w1×G + w2×Y - w3×(R + N), where w1, w2, and w3 represent weight coefficients, and w1 + w2 + w3 = 1.
[0085] The social relationship analysis module includes a social relationship graph construction unit and a social complexity analysis unit;
[0086] The social relationship graph construction unit constructs a target user set U according to the target user numbers corresponding to the target user behavior data, and U = {u1, u2,..., um}, where uj represents the j-th target user number, j ranges from 1 to m, and m represents the number of target users; for each target user, the corresponding social relationship data is obtained to construct a social relationship graph S(V, E), where V represents nodes, that is, the relevant users XU extracted from the social relationship data corresponding to the target user, which can include friend lists, followed users, comment interactions, like interactions, etc.; E represents edges. If there is social relationship data between two users, an edge is established in the social relationship graph, and the edge only indicates the existence of a social relationship, and the length of the edge has no meaning.
[0087] The social complexity analysis unit extracts all social paths for the social relationship graph S(V, E) of each target user to obtain the social path set Pa of the corresponding target user; the social path represents a path starting from the target user to any other relevant user; for each target user of the short video to be analyzed, in combination with the corresponding social path set Pa, the social complexity Z of the target user is analyzed. The specific calculation formula is: Z = [1 / N(Pa)]·Σpk∈PaL(pk), where N(Pa) represents the number of social paths in the social path set Pa, pk represents the k-th social path, and L(pk) represents the number of users in the k-th social path.
[0088] The abnormal user identification and review module includes an abnormal user identification unit and a review unit;
[0089] The abnormal user identification unit comprehensively analyzes based on the singularity of the target user's behavior pattern, content compliance, and social relationship complexity to obtain an abnormal user comprehensive score H, and H = a×C + b×(1 - F) + c×(1 - Z); compare the abnormal user comprehensive score H of the target user of the short video to be analyzed with the threshold H0. If H > H0, the target user is marked as an abnormal user; where the threshold H0 is obtained through the same analysis process as the target user based on the standard users in the database, and the minimum value of the abnormal user comprehensive scores corresponding to the standard users is selected as the threshold H0;
[0090] The review unit obtains the corresponding abnormal user numbers according to the identification results of the abnormal users, outputs the abnormal user numbers to relevant personnel for review, and screens out the abnormal traffic data of the short video to be analyzed according to the review results.
[0091] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0092] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A short video traffic anomaly analysis system based on big data analysis, characterized by: The system includes: a data collection and preliminary screening module, a user behavior analysis module, a comment content analysis module, a social relationship analysis module, and an abnormal user identification and review module; The data collection and preliminary screening module obtains the traffic data and user behavior data of the short video to be analyzed within the selected time period, and screens the traffic data to be analyzed and extracts the target user number and the corresponding target user behavior data according to the preset rules or thresholds; The user behavior analysis module analyzes the behavior pattern of the target user in a selected time period based on the target user behavior data, and performs a single evaluation on the behavior pattern of the target user based on the behavior pattern analysis result; The comment content analysis module extracts the comment information of the target user in the short video to be analyzed, performs keyword extraction and semantic analysis on the comment information and the short video to be analyzed, and evaluates the consistency between the comment information of the target user and the content of the short video to be analyzed; The social relationship analysis module constructs a target user set according to the target user number corresponding to the target user behavior data; obtains the corresponding social relationship data for each target user, thereby constructing a social relationship graph of the target user; and analyzes the social complexity of the target user of the short video to be analyzed based on the social relationship graph; The abnormal user identification and review module identifies abnormal users based on the analysis results of the target user's behavior pattern uniformity, content conformity and correlation degree; outputs the abnormal user number to relevant personnel for review, and screens out abnormal traffic data based on the review results.
2. According to the short video traffic anomaly analysis system based on big data analysis according to claim 1, it is characterized by: The data collection and preliminary screening module includes a data collection unit and a data preliminary screening unit; The data collection unit is used to collect traffic data and user behavior data of the short video to be analyzed within a selected time period, wherein the traffic data refers to statistical information of various interactions and viewing behaviors of the short video to be analyzed; The user behavior data refers to various behavior information records of users in the process of watching the short video to be analyzed; The data preliminary screening unit preliminarily screens the traffic data to be analyzed of the short video to be analyzed within the selected time period according to a preset rule or threshold, and extracts the corresponding target user number and target user behavior data based on the traffic data to be analyzed; The traffic data to be analyzed refers to traffic data that is initially determined to be normal under a preset rule or threshold.
3. According to claim 1, a short video traffic anomaly analysis system based on big data analysis is characterized in that: The user behavior analysis module includes a standard user data analysis unit, a behavior pattern recognition unit, and a behavior pattern uniqueness evaluation unit; The standard user data analysis unit extracts standard user behavior data from the database and pre-processes the standard user behavior data; Analyze the preprocessed standard user behavior data to obtain a list of standard user behavior patterns; The behavior pattern recognition unit pre-processes the target user behavior data of the target user in the selected time period, and analyzes the pre-processed target user behavior data with the standard user behavior data in the behavior pattern list of the standard user, thereby identifying the behavior pattern of the target user; The behavior pattern uniformity evaluation unit calculates the behavior pattern uniformity score of the target user within the selected time period according to the analysis result of the behavior pattern of the target user.
4. According to claim 3, a short video traffic anomaly analysis system based on big data analysis is characterized in that: The standard user data analysis unit obtains a list of standard user behavior patterns. The specific analysis process is as follows: Obtain standard user behavior data of a selected period from a database, wherein the standard user refers to a user who has been checked by relevant personnel and has no abnormal behavior; extract time series data from the standard user behavior data and perform standardization processing to obtain a behavior data sequence B(t), and B(t) = {b1(t), b2(t), ..., bn(t)}, wherein B(t) represents the behavior data sequence of the standard user at time t, b1(t) represents the intensity or frequency of the first behavior, b2(t) represents the intensity or frequency of the second behavior, and so on, bn(t) represents the intensity or frequency of the nth behavior; n represents the number of types of all behaviors of the standard user; calculate the similarity between the behavior data sequences at different times, and according to the similarity between the behavior data sequences at different times, use a clustering algorithm to divide the behavior data sequences at different times into different behavior patterns, thereby obtaining a behavior pattern list BD of the standard user, and each behavior pattern p corresponds to the behavior data sequences of several standard users.
5. According to claim 4, a short video traffic anomaly analysis system based on big data analysis is characterized in that: The specific content of the behavior pattern recognition unit recognizing the behavior pattern of the target user in the selected time period includes: Obtain the target user's behavior data within the selected time period, convert the target user's behavior data into time series data, and perform standardization processing to obtain the behavior data sequence X(t), and X(t) = {x1(t), x2(t), ..., xn(t)}, where X(t) represents the target user's behavior data sequence at time t, xi(t) represents the intensity or frequency of the i-th behavior; t ranges from 1 to n, and n represents the number of behavior types; The target user's behavior data sequence X(t) at different moments is sequentially similarly calculated with the behavior data sequence B(t) of each behavior pattern p in the standard user's behavior pattern list BD, and the average value is calculated, so as to obtain the similarity Sp between the target user's behavior data sequence X(t) at different moments and the behavior data sequence B(t) corresponding to each behavior pattern p in the standard user's behavior pattern list BD; the behavior pattern p with the largest similarity Sp is selected as the behavior pattern of the target user's behavior data sequence X(t), and all the behavior data sequences X(t) of the target user in the selected time period are traversed, so as to obtain the behavior pattern of the target user in the selected time period.
6. According to claim 5, a short video traffic anomaly analysis system based on big data analysis is characterized in that: The behavior pattern uniformity evaluation unit calculates the behavior pattern uniformity score of the target user in the selected time period, and the specific calculation process is as follows: In chronological order, the behavior patterns of the behavior data sequence X(t) at different times are summarized, and the number of types of the behavior patterns of the target user in the selected time period is counted, and the number of types of the behavior patterns of the target user in the selected time period is used as the behavior consistency measurement parameter M; Compare the deviations of the behavior patterns of two adjacent behavior data sequences X(t) in sequence, traverse the behavior data sequences X(t) of all target users in the selected time period, calculate the average deviation, and thus obtain the change measurement parameter D of the behavior pattern of the target user in the selected time period; According to the behavior consistency measurement parameter M and change measurement parameter D of the target user in the selected time period, the behavior pattern uniformity score C of the target user in the selected time period is calculated, and C = k / (α×M+β×D), where k represents the adjustment parameter, α and β represent weight coefficients, and α+β=1.
7. According to claim 1, a short video traffic anomaly analysis system based on big data analysis is characterized in that: The review content analysis module includes a text preprocessing and keyword analysis unit and a semantic analysis and conformity assessment unit; The text preprocessing and keyword analysis unit extracts the comment information of the target user under the short video to be analyzed, performs text preprocessing on the comment information of the target user under the short video to be analyzed, thereby obtaining a text set W of the target user's comment information; performs format conversion on the short video to be analyzed, converts the video data into text data, thereby obtaining a corresponding text set V; for the text set W of the target user's comment information, counts the proportion R of repeated comments and the proportion N of meaningless characters; extracts the keywords with the highest TF-IDF weights from the text set W and the text set V, respectively, and records them as the keyword set Kw and the keyword set Kv, respectively; calculates the keyword matching degree G of the keyword set Kw and the keyword set Kv, and the specific calculation formula is: G=|Kw∩Kv| / |Kw∪Kv|; The semantic analysis and conformity assessment unit maps the content of the target user's comment information and the content of the short video to be analyzed into semantic vectors through a pre-trained NLP model, which are expressed as: Ew and Ev; according to the semantic vectors Ew and Ev, the semantic correlation Y between the two is calculated, and the specific calculation formula is: Y=(Ew·Ev) / ||Ew||||Ev||; combined with the proportion of repeated comments R and the proportion of meaningless characters N, the keyword matching degree G and the semantic correlation degree Y, the target user's comment information and the content conformity F of the short video to be analyzed are comprehensively evaluated, and F=w1×G+w2×Y-w3×(R+N), wherein w1, w2 and w3 represent weight coefficients, and w1+w2+w3=1.
8. According to claim 1, a short video traffic anomaly analysis system based on big data analysis is characterized in that: The social relationship analysis module includes a social relationship graph construction unit and a social complexity analysis unit; The social relationship graph construction unit constructs a target user set U according to the target user number corresponding to the target user behavior data, and U={u1,u2,...,um}, where uj represents the jth target user number, j ranges from 1 to m, and m represents the number of target users; for each target user, the corresponding social relationship data is obtained, and a social relationship graph S(V,E) is constructed, where V represents a node, that is, a related user XU extracted from the social relationship data corresponding to the target user, which may include a friend list, a followed user, a comment interaction, a like interaction, etc.; E represents an edge. If there is social relationship data between two users, an edge is established in the social relationship graph, and the edge only represents the existence of a social relationship, and the length of the edge is meaningless; The social complexity analysis unit extracts all social paths from the social relationship graph S(V,E) of each target user, thereby obtaining the social path set Pa of the corresponding target user; the social path represents the path from the target user to any other related user; for each target user of the short video to be analyzed, the social complexity Z of the target user is analyzed in combination with the corresponding social path set Pa, and the specific calculation formula is: Z=[1 / N(Pa)]·Σpk∈PaL(pk), where N(Pa) represents the number of social paths in the social path set Pa, pk represents the kth social path, and L(pk) represents the number of users in the kth social path.
9. According to claim 1, a short video traffic anomaly analysis system based on big data analysis is characterized in that: The abnormal user identification and audit module includes an abnormal user identification unit and an audit unit; The abnormal user identification unit comprehensively analyzes the target user's behavior pattern uniformity, content conformity and social relationship complexity, thereby obtaining an abnormal user comprehensive score H, where H=a×C+b×(1-F)+c×(1-Z); compares the abnormal user comprehensive score H of the target user of the short video to be analyzed with the threshold H0, and if H>H0, the target user is marked as an abnormal user; wherein the threshold H0 is obtained by performing the same analysis process on the target user according to the standard users in the database, and the minimum value of the abnormal user comprehensive score corresponding to the standard users is selected as the threshold H0; The review unit obtains the number of the corresponding abnormal user according to the identification result of the abnormal user, outputs the abnormal user number to relevant personnel, and the relevant personnel review it, and screens out the abnormal traffic data of the short video to be analyzed according to the review result.
Citation Information
Cited By
Method for synchronizing and publishing works of multiple types of social media accounts
CN121283991A
A multi-source data fusion abnormal pattern recognition method and system
CN122554664A