An internet of things card operator data monitoring method and system based on a clustering algorithm
By combining clustering algorithms and multi-dimensional feature analysis with HDBSCAN clustering and machine learning models, the problem of system misjudgment caused by data deviations of IoT card operators was solved, achieving efficient anomaly detection and equipment fault identification, and ensuring the stable operation of IoT cards.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG SIJI TECH CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-24
AI Technical Summary
Discrepancies exist in the data from IoT card operators in the power industry, leading to system misjudgments of device offline status, resource waste, and equipment operational risks. Furthermore, data errors from different operators affect unified management and overall monitoring effectiveness.
An IoT card operator data monitoring method based on clustering algorithm is adopted. The HDBSCAN clustering method is used for multi-density clustering. Combined with a multi-dimensional evaluation index system and dynamic weight adjustment, anomaly detection and equipment fault identification are performed. Fraud detection is carried out using machine learning model and a closed-loop verification system is established.
It improves the accuracy of anomaly detection, reduces false alarms and false negatives, realizes intelligent equipment failure and fraud detection, and ensures the stable operation of IoT cards.
Smart Images

Figure CN122451501A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of IoT card technology in the power industry, and in particular to an IoT card operator data monitoring method and system based on clustering algorithm. Background Technology
[0002] As a communication carrier for critical infrastructure (such as electricity consumption information collection and transmission line channel visualization), IoT cards in the power industry can be problematic if there are deviations in the operator's data (such as incorrect traffic statistics or false status reports). This could lead to the system misjudging the device as offline, affecting statistical results and even fault warnings and maintenance responses. Failure to promptly alert for excessive traffic can result in significant resource waste. Since the data provided by the operator is upstream data, the inability to capture errors in a timely manner can lead to the accumulation and deepening of problems. Small data deviations may go undetected and gradually accumulate into major issues, ultimately affecting the actual operation of the equipment and even posing a risk of accidents.
[0003] Data errors from different operators can disrupt unified management, leading to inconsistent processing strategies and further complicating the problem. For example, simultaneous data delays from mobile operators and formatting errors from telecom operators may prevent the system from correctly integrating information, impacting overall monitoring. Furthermore, in-depth analysis of how data errors trigger a chain reaction becomes an extremely labor-intensive task with minimal returns and a very low cost-effectiveness ratio.
[0004] Clustering algorithms can divide data into different groups, making them suitable for detecting abnormal patterns. When monitoring data accuracy, clustering can identify outlier data points that do not conform to the majority of data patterns, potentially indicating inaccurate data. For example, similar devices in the same region should have similar traffic usage patterns. If a device's data deviates significantly from the cluster center, it may mean that data reporting is incorrect or there is a problem with the operator's data. Clustering algorithms can automatically and efficiently process massive amounts of data without the need for predefined rules, and can quickly detect anomalies. Summary of the Invention
[0005] To address the aforementioned problems, this invention provides a method and system for monitoring IoT card operator data based on clustering algorithms.
[0006] In a first aspect, the present invention provides an IoT card operator data monitoring method based on a clustering algorithm, which adopts the following technical solution: A method for monitoring IoT SIM card operator data based on clustering algorithms, comprising: Obtain IoT SIM card operator data; Preprocess the acquired IoT card operator data; A multi-dimensional evaluation index system was constructed, and a dynamic weight fusion model was used to dynamically adjust the weights of the preprocessed IoT card operator data. Time-series features were constructed from the adjusted IoT SIM card operator data; Anomaly detection of feature data is performed using the HDBSCAN clustering method. Specifically, multi-density clustering of IoT card operator data is performed using the HDBSCAN clustering algorithm, and anomaly detection is performed based on multi-density clustering. Perform fraud detection and device malfunction identification on IoT cards that show abnormalities.
[0007] Furthermore, the data preprocessing of the acquired IoT card operator data includes data standardization based on a multi-source heterogeneous data adaptive cleaning framework and a triple mapping mechanism. Specifically, numerical data, character data, and date data are converted through data type mapping, and the field value ranges in different data sources are linearly mapped to a unified standard value range. The field names in different data sources are semantically mapped after lexical analysis.
[0008] Furthermore, the construction of a multi-dimensional evaluation index system involves dynamically adjusting the weights of the preprocessed IoT card operator data using a dynamic weight fusion model. This includes calculating the field missing rate, value range compliance rate, and temporal contradiction law of the IoT card operator data, constructing a judgment matrix, and using an improved model entropy weight method while dynamically adjusting the weights by introducing a data fluctuation factor.
[0009] Furthermore, the construction of time-series features for the adjusted IoT card operator data includes decomposing the IoT card operator data using a time series decomposition method, performing data feature statistics on the decomposed data based on a sliding window, and introducing a recurrent neural network for time-series feature extraction and learning.
[0010] Furthermore, the multi-density clustering of IoT card operator data using the HDBSCAN clustering algorithm includes performing multi-density clustering based on the feature distribution of IoT card operator data using the HDBSCAN clustering algorithm. Specifically, the HDBSCAN clustering algorithm is used to traverse data points, and the clustering is repeated multiple times by dynamically adjusting the clustering density threshold to obtain the clustering results.
[0011] Furthermore, the anomaly detection based on multi-density clustering includes identifying low-density regions based on the multi-density clustering results to define anomalies, determining outliers through outlier analysis, detecting cross-cluster anomalies through inter-cluster difference analysis to detect abnormal usage patterns of IoT cards, and finally calculating density reachability path differences of data points based on density reachability analysis to determine anomalies based on density reachability path differences.
[0012] Furthermore, the process of performing fraud detection and device fault identification on IoT cards that detect anomalies includes multi-dimensional feature extraction and quantification from time, space and behavior dimensions, fraud probability prediction using machine learning models, device fault identification based on a rule engine, and finally fraud response and fault response based on a linkage handling mechanism.
[0013] Secondly, an IoT SIM card operator data monitoring system based on clustering algorithms includes: The data acquisition module is configured to acquire data from the IoT card operator; The preprocessing module is configured to preprocess the acquired IoT card operator data. The weight adjustment module is configured to construct a multi-dimensional evaluation index system and use a dynamic weight fusion model to dynamically adjust the weights of the preprocessed IoT card operator data. The feature construction module is configured to construct time-series features from the adjusted IoT SIM card operator data. The anomaly detection module is configured to use the HDBSCAN clustering method to perform anomaly detection on feature data. Specifically, the HDBSCAN clustering algorithm is used to perform multi-density clustering on IoT card operator data, and anomaly detection is performed based on multi-density clustering. The discrimination module is configured to perform fraud detection and device malfunction identification on IoT cards that are found to be abnormal.
[0014] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the IoT card operator data monitoring method based on a clustering algorithm.
[0015] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor as described in the IoT card operator data monitoring method based on clustering algorithm.
[0016] In summary, the present invention has the following beneficial technical effects: 1. By extracting and quantifying multi-dimensional features, and combining feature analysis of time, space, behavior and other dimensions, specific algorithms and formulas are used to accurately calculate the abnormal indicators of each dimension. For example, the Haversine formula is used to calculate geographical distance to judge spatial anomalies, and the flow change rate formula is used to judge flow anomalies. This greatly improves the accuracy of anomaly detection, reduces false alarms and false negatives, and can promptly discover potential problems in the use of IoT cards.
[0017] 2. A combination of rule engines and machine learning models is used for fraud detection and equipment fault identification. The rule engine sets clear trigger conditions for rapid response to common anomalies; machine learning models, such as logistic regression models, can learn complex patterns from large amounts of data to accurately predict fraud probabilities. In equipment fault identification, methods such as cosine similarity are used to determine fault recovery status, achieving an intelligent and efficient decision-making process and improving problem-solving efficiency.
[0018] 3. The established closed-loop verification system retains signaling data for anomaly tracing, conducts A / B testing to optimize decision tree rules, and regularly performs red-blue team exercises to verify the accuracy of discrimination, continuously optimizing and improving the entire detection and discrimination system. This enables the system to adapt to constantly changing network environments and equipment conditions, continuously improving its ability to detect and handle IoT card anomalies, and ensuring the stable and secure operation of IoT cards. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the method in Embodiment 1 of the present invention. Detailed Implementation
[0020] The present invention will be further described in detail below with reference to the accompanying drawings.
[0021] Example 1 Reference Figure 1 This embodiment of an IoT SIM card operator data monitoring method based on clustering algorithm includes: Obtain IoT SIM card operator data; Preprocess the acquired IoT card operator data; A multi-dimensional evaluation index system was constructed, and a dynamic weight fusion model was used to dynamically adjust the weights of the preprocessed IoT card operator data. Time-series features were constructed from the adjusted IoT SIM card operator data; Anomaly detection of feature data is performed using the HDBSCAN clustering method. Specifically, multi-density clustering of IoT card operator data is performed using the HDBSCAN clustering algorithm, and anomaly detection is performed based on multi-density clustering. Perform fraud detection and device malfunction identification on IoT cards that show abnormalities.
[0022] Specifically, it includes the following steps: S1. Acquire IoT SIM card operator data. This involves interfacing with various IoT SIM card operators to collect relevant data in real-time or periodically, including data usage, signal strength, device online status, and data plan information. To ensure data comprehensiveness, a multi-channel data collection mechanism is established, acquiring data not only from the operators' core databases but also from their edge node data caches to capture rapidly changing data.
[0023] S2. Data preprocessing is performed on the acquired IoT SIM card operator data. A multi-source heterogeneous data adaptive cleaning framework is adopted, and data standardization is achieved based on a triple mapping mechanism. Determine the data type for each field and calculate the percentage of each data type across all fields. Determine the value range for each field and record its minimum and maximum values. Check for missing values and calculate the percentage of missing values in each field.
[0024] Data type mapping Identifying data types: For numeric data, the criteria are whether the field value consists only of numbers (including integers and decimals) and whether it conforms to common numerical arithmetic rules. For character data, check whether the field value is a string composed of letters, numbers, special characters, etc. For date data, determine its type based on common date formats (such as YYYY-MM-DD, DD / MM / YYYY, etc.).
[0025] Type Conversion: Numeric data is uniformly converted to floating-point number type. Date data is uniformly converted to standard date format (e.g., YYYY-MM-DD). Conversion Formula: Assuming the original date format is DD / MM / YYYY, when converting to the standard format, the day, month, and year information are extracted and recombine in the order YYYY-MM-DD.
[0026] Range mapping Determine the value range: For each field, calculate its minimum value in different data sources. and maximum value ,in Indicates the first One data source.
[0027] Linear mapping: The goal is to map the field value ranges from different data sources to a unified standard value range. For any data value in the data source, use the following linear mapping formula for transformation: .
[0028] in, It is the mapped data value. and These are the lower and upper limits of the standard value range. These are the original data values. and These are the minimum and maximum values of the field in the data source, respectively.
[0029] Semantic mapping (1) Semantic Analysis: Perform lexical analysis on field names from different data sources to extract keywords. Calculate the semantic similarity between field names, using methods such as cosine similarity. Assume the word vectors of the two field names are respectively... and Then their cosine similarity is calculated using the following formula: in, It is the dot product of two vectors. and These are the magnitudes of the two vectors, respectively.
[0030] Determine which fields have the same or similar semantics based on semantic similarity.
[0031] Mapping rule establishment: For fields with the same semantics but different value representations, establish mapping rules. For example, if "Male" represents male in one data source and "M" represents male in another data source, then establish a mapping rule to convert "M" to "Male".
[0032] Data cleaning and validation (1) Outlier cleanup: Outliers were detected using the Z-score method. For each field, its mean was calculated. and standard deviation For the data values in the field Calculate its Z-score: Data values whose absolute Z-score is greater than a certain threshold (usually 3) are considered outliers and are deleted or corrected.
[0033] (2) Verify data quality: Calculate the completeness of the data, i.e., the proportion of non-missing values: The accuracy of the data can be checked by comparing it with known correct data.
[0034] Verify data consistency by checking whether the values of the same semantic field are consistent across different data sources.
[0035] S3. Construct a multi-dimensional evaluation index system, and use a dynamic weight fusion model to dynamically adjust the weights of the preprocessed IoT card operator data. This includes calculating the field missing rate, value range compliance rate, and temporal contradiction law of the IoT card operator data, constructing a judgment matrix, and using an improved entropy weight method by introducing a data fluctuation factor to dynamically adjust the weights. (1) Calculate the field missing rate Counting Missing Values: For each field, iterate through all data records under that field and count the number of missing values. Let the first field be the first... The number of missing values in each field is The total number of records is .
[0036] Calculate the field missing rate: Calculate the missing rate for each field according to the formula. The formula is: (2) Calculate the compliance rate of the value range Determine the value range: For each field, based on business knowledge or historical data, determine its reasonable value range, let the first field be... The minimum value of each field is The maximum value is .
[0037] Check data compliance: Iterate through all data records under each field and determine whether the data values are within a reasonable range. Let the number of data records with reasonable values in the i-th field be . .
[0038] Calculate the range compliance rate: Calculate the range compliance rate for each field according to the formula. The formula is: (3) Calculate the temporal contradiction law Extracting time-series data: Extract time-related fields and corresponding business data fields from IoT card operator data to form time-series data pairs.
[0039] Define time sequence logic rules: Based on business logic, define reasonable time sequence rules. For example, for traffic usage data, the cumulative traffic at a later time point should be greater than or equal to the cumulative traffic at the previous time point.
[0040] Check for temporal discrepancies: Iterate through the time series data pairs in chronological order and check for any violations of temporal logic rules. Let the number of data pairs violating temporal logic rules in the i-th field be . .
[0041] Calculate the timing contradiction law: Calculate the timing contradiction law for each field according to the formula. The formula is: Because the number of time series data pairs is one less than the total number of records, it is represented by N-1.
[0042] (4) Construct the judgment matrix Determine the evaluation metrics: Include the field missing rate. Value range compliance rate Contradictory Law of Time As an indicator for evaluating the data quality of IoT card operators.
[0043] Pairwise comparison of indicator importance: Using methods such as expert scoring or the Analytic Hierarchy Process (AHP), the three indicators are compared pairwise to determine their relative importance. Assumptions Indicates the first The importance judgment value of each indicator relative to the i-th indicator, when hour, ;when When, the value is determined based on the importance comparison result, such as This indicates the importance of the field missing rate relative to the value range compliance rate. If the field missing rate is slightly more important than the value range compliance rate, then... ,on the contrary .
[0044] Constructing the judgment matrix: Based on the comparison results above, construct a 3x3 judgment matrix: An improved model using entropy weighting and incorporating a data fluctuation factor to dynamically adjust the weights. Data standardization: Field missing rate Value range compliance rate Contradictory Law of Time Standardization is performed to eliminate the influence of dimensions. A linear transformation method is used to map the data to... Range. For indicators where a larger value is generally better (such as range compliance rate), the standardized formula is: For metrics where smaller values are better (such as field missing rate and temporal inconsistency law), the standardized formula is: ,in It is the first The first sample Individual indicator values, It is the standardized value.
[0045] Calculate information entropy: Based on the standardized data, calculate the information entropy of each indicator. The formula is: ,in , It is the sample size. .
[0046] Introducing data volatility factors: Calculating the data volatility factor for each indicator. The data volatility factor can be obtained by calculating the standard deviation of the indicator values, i.e. The greater the fluctuation in the data, the more important the information contained in the indicator, and the greater its influence should be given in subsequent weight calculations.
[0047] Calculate the weights: Combine information entropy and data fluctuation factor to calculate the weight of each indicator. First, calculate the coefficient of difference. Then calculate the overall weight: ,in It refers to the number of indicators.
[0048] S4. Construct time-series features for the adjusted IoT card operator data, including decomposing the IoT card operator data using time series decomposition methods, performing data feature statistics on the decomposed data based on a sliding window, and introducing a recurrent neural network for time-series feature extraction and learning.
[0049] in, (1) Time series decomposition Applying STL decomposition: Initialize parameters, set the Loess smoothing parameter and window width. The window width determines the number of neighboring data points considered when performing local weighted regression.
[0050] Calculate the trend term : By analyzing the original time series Applying Loess smoothing removes the effects of seasonality and random fluctuations, yielding the trend term. The specific process involves performing multiple locally weighted regressions on the time series to gradually fit the trend.
[0051] Calculate the seasonal term From the original time series Subtract the trend term This yields a sequence containing seasonal and random components. Then, periodic analysis was performed on the sequence, and the seasonal term was extracted by calculating the average value of different periods. .
[0052] Calculate random terms : Using the original time series Subtract trend item and seasonal items ,Right now , thus obtaining a random item.
[0053] (2) Data feature statistics based on sliding window Define sliding window parameters: Set the size of the sliding window. (i.e., the number of data points contained within the window) and sliding step size (The distance the window moves each time).
[0054] Traversing the time series: Starting from the beginning of the time series, move the window according to the sliding step size.
[0055] Calculate window features: For the data within each window (including trend, seasonal, or random data), calculate the following statistical features: Mean: ,in These are data points within the window. It is the index of the window's starting position.
[0056] variance: Maximum value: Minimum value: Median: The value at the middle position after sorting the data points within the window (if...). (if odd) or the average of the two middle values (if) (If it is an even number).
[0057] Storing feature values: The calculated statistical feature values of each window are stored to form a new feature sequence.
[0058] (3) Introduce recurrent neural networks for temporal feature extraction and learning. The feature sequences calculated using a sliding window are used as input data for the GRU recurrent neural network. The data is then normalized using the min-max normalization method. ,in and These are the minimum and maximum values of the feature sequence, respectively.
[0059] Constructing a recurrent neural network model: Choose the network architecture: Determine the number of GRU layers, the number of hidden units, and other parameters. Construct a network with two LSTM layers, each with 64 hidden units.
[0060] Input and output settings: The input dimension of the input layer is the dimension of the feature sequence (i.e., the number of features input each time, such as the number of features calculated earlier, such as the mean, variance, etc.). The output layer is set according to the specific task. When predicting the value at the next time point, the output layer dimension is 1.
[0061] Training the model: Define the loss function: Mean Squared Error (MSE) Loss Function: ,in These are model predictions. It is the actual value. It refers to the number of samples.
[0062] Select the Adam optimizer and set parameters such as the learning rate. Train the model by inputting the prepared input data and corresponding target data, and updating the model parameters using the backpropagation algorithm to minimize the loss function.
[0063] Feature extraction and learning: During training, the recurrent neural network automatically learns the temporal features in the input data. After training, the output of the model's hidden layers can be used as the extracted temporal features, which contain long-term dependencies and complex pattern information of the time series.
[0064] S5. Perform multi-density clustering on IoT SIM card operator data using the HDBSCAN clustering algorithm, including multi-density clustering based on the characteristic distribution of IoT SIM card operator data using the HDBSCAN clustering algorithm, and dynamically adjust the clustering density threshold. in, (1) Initialize HDBSCAN parameters Set an initial density threshold: Based on preliminary data analysis and experience, set an initial density threshold. This is used to define the neighborhood radius of a core point. A core point is a point within a certain radius of its neighborhood. The neighborhood contains at least A point.
[0065] Determine the minimum number of points Choose an appropriate value , The choice of [aspect name] affects the clustering results and the identification of noise points. In this embodiment, The value is related to the dimension and density of the data. ≥ (data dimension + 1).
[0066] (2) Execute the HDBSCAN clustering algorithm Traversing data points: Starting from the first point in the dataset, traverse each data point in turn. .
[0067] Calculate neighborhood points: for each data point Calculate its in All points within the neighborhood are denoted as .
[0068] Key point to judge: If Then point The core point is to create a new cluster. And add points in its neighborhood to the cluster. Then, new core points are searched from these neighborhood points, and their neighborhood points are also added to the cluster. The process continues until no new points can be added to the cluster.
[0069] Mark noise points: if Then point These are noise points and are marked as unclassified.
[0070] (3) Dynamically adjust the cluster density threshold Evaluate clustering results: Evaluate the current clustering results, such as calculating the silhouette coefficient of the clusters. The silhouette coefficient is used to measure the tightness and separation of clusters, and its calculation formula is as follows: ,in It is the average distance from a point to other points in the same cluster. It is a point The average distance to the nearest point in other clusters. The silhouette coefficient of the entire clustering result is the average of the silhouette coefficients of all points.
[0071] Adjust density threshold: Dynamically adjust the density threshold based on the evaluation results. If the silhouette coefficient is low, it indicates that the clustering effect is not ideal, which may be due to an improper density threshold setting. In this case, you can try decreasing or increasing the density threshold. The specific adjustment methods are as follows: If the contour coefficient is less than a certain preset threshold Furthermore, the current number of clusters is relatively large, indicating that the density threshold may be too small, leading to over-clustering of the data. In this case, increasing the density threshold is appropriate. ,in It is an adjustment step size, which is 0.1 in this embodiment.
[0072] If the contour coefficient is less than Furthermore, the current number of clusters is relatively small, indicating that the density threshold may be too high, causing data to be merged into only a few clusters. In this case, reducing the density threshold is appropriate. .
[0073] Re-clustering: using the adjusted density threshold Then, re-execute the HDBSCAN clustering algorithm and repeat step 3 until the silhouette coefficient of the clustering result reaches a satisfactory level or the preset number of iterations is reached.
[0074] S6. Anomaly detection based on multi-density clustering, wherein, (1) Identification of low-density areas Identifying low-density regions: The HDBSCAN algorithm categorizes data points into core points, boundary points, and noise points. Core points are located in high-density regions, boundary points are at the edges of high-density regions, and noise points belong to low-density regions. In the clustering results, data points marked as noise points are highly likely to be located in low-density regions.
[0075] Defining outliers: Based on the density-based anomaly detection hypothesis, data points in low-density areas have significantly different distribution patterns from their surrounding data points, so these low-density data points can be preliminarily identified as outliers.
[0076] By traversing the clustering results, all data points marked as noise points are filtered out, and these points are the candidate set of outliers based on low-density region determination.
[0077] (2) Outlier analysis Calculating outlier severity: For each cluster, calculate the distance from each data point to the cluster center. Use the Euclidean distance formula, assuming the cluster size is... have Data points Cluster center is data points The Euclidean distance formula to the cluster center is: in It is the dimension of the data. Data points The Values of each dimension Cluster center The The values of each dimension.
[0078] Set an outlier threshold: Based on the data distribution and business requirements, set an outlier threshold. If the distance from a data point to its cluster center is greater than a threshold... If the data point is considered an outlier, it is considered an anomaly.
[0079] Algorithm steps: For each cluster First, calculate its cluster centers. , ,in It is clustering The number of data points in the data.
[0080] Clustering Each data point in Calculate according to the above Euclidean distance formula .
[0081] Will Outlier threshold In comparison, if Then Mark as an anomaly.
[0082] 3. Inter-cluster difference analysis Compare clustering characteristics: Different clusters represent different data patterns. Analyze the characteristics of each cluster, such as mean, variance, and data point distribution range. For IoT SIM card traffic usage data clustering, different clusters may represent IoT SIM cards in different usage scenarios, such as low-traffic monitoring device cards and high-traffic video transmission device cards.
[0083] Detecting cross-cluster anomalies: If a data point's features differ significantly from the features of its assigned cluster, but are more similar to features of other clusters, then this data point may be an outlier. Cosine similarity is used to calculate the similarity between a data point and different cluster features. Assume the data point... Clustering and The feature vectors are respectively , , The cosine similarity formula is: like much smaller ,and I was assigned to ,but It might be an outlier.
[0084] Algorithm steps: Extract each cluster eigenvectors It can be the mean vector of a cluster, etc.
[0085] For each data point Calculate its eigenvector .
[0086] For each data point Calculate its relationship with its cluster. )( express (The cluster number it belongs to) and all other clusters cosine similarity and .
[0087] Set similarity difference threshold ,like: Then Mark as an anomaly.
[0088] 4. Based on density reachability analysis Tracing density reachable paths: In HDBSCAN clustering, core points are connected by density reachability relationships. For each data point, its density reachable path is traced. If a data point's density reachable path is very short or differs significantly from the density reachable paths of other normal data points, it may indicate that the data point is abnormal.
[0089] Anomaly Path Identification: Normal data points are typically connected to larger clusters via reasonable density-reachable paths. If a data point can only be connected to other points via a few density-reachable steps, or if the points it connects to are different from those connected to most normal data points, then this data point may be an anomaly.
[0090] in, Construct a density-reachable graph, where nodes are core points and edges represent density-reachable relationships.
[0091] For each data point ,like It's the core point, from Start by performing a breadth-first search (BFS) or depth-first search (DFS) and record the density-reachable path length. and the set of nodes on the path .
[0092] Calculate the average path length of all normal data points (non-anomaly candidate points). And the characteristics of the average path node set.
[0093] Set path length threshold Difference threshold between path nodes ,like or The difference metric between the path node set and the normal data point set is greater than Then Mark as an anomaly.
[0094] S7. Perform fraud detection and device malfunction identification on IoT cards with abnormal detection, among which, (1) Multidimensional feature extraction and quantization Time dimension: Calculate the deviation between the time of an anomaly occurrence and the historical average usage time of the equipment. The formula is ,in The time when the anomaly occurred. This is the historical average usage time. If... Greater than the set time threshold If so, it is marked as a time dimension anomaly.
[0095] Spatial dimension: The geographical distance between the anomaly location and the device registration location is calculated using the Haversine formula. Let the latitude and longitude of the registered location be... The latitude and longitude of the abnormal location are Earth's radius is ,but: like Greater than the set distance threshold If so, it is determined to be an anomaly in spatial dimension.
[0096] Behavioral dimension: Flow characteristics: Calculating the rate of change of flow The formula is ,in For current traffic, This is the historical average flow rate. If... Greater than the flow rate change threshold If so, it is marked as an abnormal traffic flow.
[0097] Connection characteristics: The number of connections per unit time. ,like Greater than the connection count threshold If so, it is determined to be a connection error.
[0098] (2) Intelligent classification decision Fraud detection: Rule Engine: Set rules such as high-frequency activation in different locations (number of activations) Greater than the threshold for the number of times an activity can be activated in a different location And abnormal traffic A fraud alert is triggered when the number of users accessing the site exceeds the threshold and an unauthorized APN (APN not in the authorized list) is used.
[0099] Machine learning model: A logistic regression model is used for fraud probability prediction. Let the feature vector be... Including the extracted multi-dimensional features such as time, space, and behavior, the model formula is: ,in This indicates fraud. These are the model parameters, obtained through training on a large amount of labeled data. When predicting probabilities... Greater than the set fraud probability threshold At that time, it was determined to be fraud.
[0100] External verification: Call the operator's risk control system interface to query whether the SIM card has abnormal status information such as being reported lost or marked as a cloned card.
[0101] Equipment fault diagnosis: Rule engine: When signal strength Less than the signal strength threshold And the number of sensor data jumps Greater than the threshold number of transitions Meanwhile, the number of heartbeat packets lost Greater than the number of lost times threshold At that time, a device fault warning is triggered.
[0102] Diagnostic logic: Execute remote command verification, such as sending a restart command and comparing data characteristics before and after the restart. Let the data feature vector before the restart be... After restarting, the data feature vector is By calculating the cosine similarity between the two: To determine the data recovery status. If the similarity is greater than the recovery similarity threshold. If so, it is determined to be a temporary fault.
[0103] Device fingerprint: Compare the IMEI / MAC with the registration information. If they do not match, it is determined that the device may have been replaced.
[0104] (3) Joint response mechanism Fraud Response: Real-time Communication Blocking: Send a blocking command to the operator to cut off the IoT SIM card's network connection using the operator's DPI (Deep Packet Inspection) policy. Generate a risk ticket and push it to the security team. The ticket includes IoT SIM card information, anomaly characteristics, and fraud determination criteria. Update the threat intelligence database, inputting the characteristic information of this fraud case for subsequent fraud detection model training and rule optimization.
[0105] Fault Response: Automatically switch to backup communication links, selecting the optimal backup communication path according to a preset link switching strategy. Push maintenance work orders to the equipment management system; the work order includes equipment information, fault characteristics, and fault determination criteria. Initiate a remote firmware upgrade process; if the fault is determined to be due to a firmware issue, send a firmware upgrade package to the device.
[0106] (4) Closed-loop verification system: Establish an abnormal event tracing mechanism and retain signaling data for a certain period of time (e.g., 30 days) so that the abnormal events can be analyzed in detail later.
[0107] The decision tree rules are continuously optimized through A / B testing. The accuracy of fraud detection and device fault identification under different rules is compared, and the optimal rule is selected.
[0108] Regularly conduct red-blue team exercises to simulate real fraud and equipment failure scenarios, verify the accuracy of the judgment, and continuously improve the system's detection and judgment capabilities.
[0109] Example 2 This embodiment provides an IoT SIM card operator data monitoring system based on a clustering algorithm, including: The data acquisition module is configured to acquire data from the IoT card operator; The preprocessing module is configured to preprocess the acquired IoT card operator data. The weight adjustment module is configured to construct a multi-dimensional evaluation index system and use a dynamic weight fusion model to dynamically adjust the weights of the preprocessed IoT card operator data. The feature construction module is configured to construct time-series features from the adjusted IoT SIM card operator data. The anomaly detection module is configured to use the HDBSCAN clustering method to perform anomaly detection on feature data. Specifically, the HDBSCAN clustering algorithm is used to perform multi-density clustering on IoT card operator data, and anomaly detection is performed based on multi-density clustering. The discrimination module is configured to perform fraud detection and device malfunction identification on IoT cards that are found to be abnormal.
[0110] A computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device, the aforementioned IoT card operator data monitoring method based on a clustering algorithm.
[0111] A terminal device includes a processor and a computer-readable storage medium, the processor being configured to implement instructions; the computer-readable storage medium being configured to store a plurality of instructions adapted for loading by the processor and executing the method thereof.
[0112] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for monitoring IoT SIM card operator data based on clustering algorithms, characterized in that, include: Obtain IoT SIM card operator data; Preprocess the acquired IoT card operator data; A multi-dimensional evaluation index system was constructed, and a dynamic weight fusion model was used to dynamically adjust the weights of the preprocessed IoT card operator data. Time-series features were constructed from the adjusted IoT SIM card operator data; Anomaly detection of feature data is performed using the HDBSCAN clustering method. Specifically, multi-density clustering of IoT card operator data is performed using the HDBSCAN clustering algorithm, and anomaly detection is performed based on multi-density clustering. Perform fraud detection and device malfunction identification on IoT cards that show abnormalities.
2. The IoT SIM card operator data monitoring method based on clustering algorithm according to claim 1, characterized in that, The data preprocessing of the acquired IoT card operator data includes data standardization based on a multi-source heterogeneous data adaptive cleaning framework and a triple mapping mechanism. Specifically, numerical data, character data, and date data are converted through data type mapping, and the field value ranges in different data sources are linearly mapped to a unified standard value range. The field names in different data sources are semantically mapped after lexical analysis.
3. The IoT SIM card operator data monitoring method based on clustering algorithm according to claim 2, characterized in that, The construction of a multi-dimensional evaluation index system utilizes a dynamic weight fusion model to dynamically adjust the weights of preprocessed IoT card operator data. This includes calculating the field missing rate, value range compliance rate, and temporal contradiction law of the IoT card operator data, constructing a judgment matrix, and using an improved entropy weight method by introducing a data fluctuation factor to dynamically adjust the weights.
4. The IoT SIM card operator data monitoring method based on clustering algorithm according to claim 3, characterized in that, The process of constructing time-series features for the adjusted IoT SIM card operator data includes decomposing the IoT SIM card operator data using a time-series decomposition method, performing data feature statistics on the decomposed data based on a sliding window, and introducing a recurrent neural network for time-series feature extraction and learning.
5. The IoT SIM card operator data monitoring method based on clustering algorithm according to claim 4, characterized in that, The method of performing multi-density clustering of IoT card operator data using the HDBSCAN clustering algorithm includes performing multi-density clustering based on the feature distribution of IoT card operator data using the HDBSCAN clustering algorithm. Specifically, the HDBSCAN clustering algorithm is used to traverse data points, and the clustering is repeated multiple times by dynamically adjusting the clustering density threshold to obtain the clustering results.
6. The IoT SIM card operator data monitoring method based on clustering algorithm according to claim 5, characterized in that, The anomaly detection based on multi-density clustering includes identifying low-density regions based on the multi-density clustering results to define anomalies, determining outliers through outlier analysis, detecting cross-cluster anomalies through inter-cluster difference analysis to detect abnormal usage patterns of IoT cards, and finally calculating the density reachability path differences of data points based on density reachability analysis to determine anomalies.
7. The IoT SIM card operator data monitoring method based on clustering algorithm according to claim 6, characterized in that, The process of fraud detection and device fault identification for abnormal IoT cards includes multi-dimensional feature extraction and quantification from time, space and behavior dimensions, calculation of the geographical distance between the abnormal location and the device registration location using the Haversine formula, fraud probability prediction using a machine learning model, device fault identification based on a rule engine, and finally fraud response and fault response based on a linkage handling mechanism.
8. An IoT SIM card operator data monitoring system based on clustering algorithm, characterized in that, include: The data acquisition module is configured to acquire data from the IoT card operator; The preprocessing module is configured to preprocess the acquired IoT card operator data; The weight adjustment module is configured to construct a multi-dimensional evaluation index system and use a dynamic weight fusion model to dynamically adjust the weights of the preprocessed IoT card operator data. The feature construction module is configured to construct time-series features from the adjusted IoT SIM card operator data; The anomaly detection module is configured to use the HDBSCAN clustering method to perform anomaly detection on feature data. Specifically, the HDBSCAN clustering algorithm is used to perform multi-density clustering on IoT card operator data, and anomaly detection is performed based on multi-density clustering. The discrimination module is configured to perform fraud detection and device malfunction identification on IoT cards that are found to be abnormal.
9. A computer-readable storage medium storing a plurality of instructions, characterized in that, The instructions are adapted to be loaded by the processor of the terminal device and executed as described in claim 1.
10. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is configured to implement instructions; and the computer-readable storage medium is configured to store multiple instructions, characterized in that, The instructions are adapted to be loaded by a processor and executed as described in claim 1.