Network Security Management System Based on Big Data
By constructing a dynamic threshold calculation model and fuzzy comprehensive evaluation, combined with machine learning optimization units, the problems of misjudgment and misjudgment of traditional account risk assessment methods are solved, and the accuracy and reliability of cybercrime account risk assessment are achieved.
Patent Information
- Application Number
- CN202510645884.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The traditional account risk assessment method relies on fixed thresholds and cannot adapt to the differences in the flow of cybercrime funds, resulting in misjudgment of telecommunications fraud and misjudgment of financial fraud, affecting the effectiveness of network security management.
A network security management system based on big data is adopted to construct a dynamic threshold calculation model through a dynamic threshold setting unit, combining the hidden Markov model, kernel density estimation and reinforcement learning algorithm, combining the fuzzy comprehensive evaluation unit to determine the membership function and weights, and using machine learning optimization unit to train the model and update it regularly to improve the accuracy of account risk assessment.
Accurately evaluate account risks, reduce misjudgments and misjudgments, improve the effectiveness and reliability of network security management, and ensure the normal development of financial services.
Smart Images

Figure CN120181987B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network security management, and in particular to a network security management system based on big data. Background Art
[0002] Network security management is a crucial technology. Account risk assessment is a crucial component in combating cybercrime. However, due to the significant differences in capital flow patterns across different cybercrime scenarios, traditional account risk assessment methods based on fixed thresholds struggle to adapt to these changes. This results in significant problems with current account risk assessments for cybercrime.
[0003] Traditional account risk assessment methods mainly rely on fixed thresholds to determine whether an account is risky. However, the capital flow of cybercrime is complex and changeable. Different types of cybercrime vary greatly in terms of transaction frequency, amount of money involved, and transaction time span. Telecom fraud cybercrime often has a high transaction frequency, relatively small amounts involved in each transaction, but the overall amount of money involved grows rapidly. In contrast, financial fraud cybercrime may have a longer transaction time span and involve large and concentrated amounts of money. Fixed thresholds cannot adapt to these differences. When dealing with telecommunications fraud, the threshold may be set too high, resulting in a large number of suspicious accounts being missed and risks cannot be detected in a timely manner. When dealing with financial fraud, normal accounts may be mistakenly identified as risky accounts due to unreasonable threshold settings, affecting the normal development of financial business. This inaccurate risk assessment method seriously restricts the effectiveness of network security management. To solve this technical problem, we provide a network security management system based on big data. Summary of the Invention
[0004] The purpose of the present invention is to provide a network security management system based on big data to solve the problems raised in the above background technology.
[0005] 1. Because traditional account risk assessment relies on fixed thresholds and cannot adapt to the differences in capital flows in cybercrime, resulting in missed detections of telecommunications fraud and misjudgments of financial fraud, this case uses a dynamic threshold setting unit to construct a dynamic threshold calculation model for different types of cybercrime. Incorporating techniques such as hidden Markov models, kernel density estimation, and reinforcement learning, the threshold is adjusted based on transaction frequency, the amount involved, and the transaction time span. The fuzzy comprehensive evaluation unit introduces fuzzy mathematical algorithms to determine the membership function, assign weights, and calculate the comprehensive account risk score. This can accurately assess account risk, reduce missed detections and misjudgments, improve the effectiveness of network security management, and ensure the normal operation of financial business.
[0006] 2. Since it is difficult to process multi-source heterogeneous data related to cybercrime in the traditional way, which affects the accuracy of risk assessment, in this case, the data acquisition and integration unit uses the DATAX multi-source heterogeneous data cleaning technology, including the preliminary cleaning, outlier detection, and data repair modules. The acquisition process has the characteristics of intelligent perception and adaptive scheduling. By using deep neural networks to analyze web page structures, ant colony algorithms to schedule tasks, and blockchain technology to cache data, high-quality data can be obtained, providing a reliable basis for subsequent risk assessment and enhancing the reliability and stability of the network security management system.
[0007] To achieve the above objectives, a network security management system based on big data is provided, including the following units:
[0008] The data acquisition and integration unit is used to collect cybercrime data from multiple data sources and process the cybercrime data using the DATAX multi-source heterogeneous data cleaning technology to obtain its fund transfer characteristics;
[0009] The dynamic threshold setting unit constructs a dynamic threshold calculation model based on the fund transfer characteristics and dynamically adjusts the threshold according to the transaction frequency, involved amount, and transaction time of the cybercrime data;
[0010] The fuzzy comprehensive evaluation unit introduces the fuzzy mathematics algorithm to determine the membership function of the cybercrime data, assigns weights to the transaction frequency, amount concentration, and account relevance respectively, calculates the comprehensive risk score of the cybercrime data accounts through fuzzy transformation, compares the comprehensive risk score with the set risk threshold, and determines the risk accounts according to the comparison results;
[0011] The machine learning optimization unit uses the historical case data to construct a cybercrime training data set, selects the neural network algorithm to train the cybercrime training data set, establishes a fund penetration threshold determination model, then regularly updates the training set with new case data, and finally determines the involved accounts in the risk accounts through the control unit.
[0012] As a further improvement of this technical solution, when the data acquisition and integration unit processes the cybercrime data using the DATAX multi-source heterogeneous data cleaning technology:
[0013] [[ID=2 | 1]] [[ID=2 | 2]]
[0014] Finally, for the abnormal data, the data repair module based on the generative adversarial network is used to repair the abnormal data.
[0015] As a further improvement of this technical solution, the acquisition process of the data acquisition and integration unit has the characteristics of intelligent perception and adaptive scheduling:
[0016] In the acquisition source detection stage, a web structure parser based on a deep neural network is used to learn and identify the data storage structure and update frequency pattern of different types of websites, and an acquisition path planning template is constructed;
[0017] When collecting from distributed data sources, a task scheduling mechanism based on the ant colony algorithm is adopted to analogize the acquisition tasks to the foraging paths of ants, and the acquisition task priorities are dynamically assigned;
[0018] And a trusted data caching mechanism based on blockchain is built in to cope with network fluctuations and data source limitations during the acquisition process, and the acquired data fragments are stored in the form of blockchain blocks.
[0019] As a further improvement of this technical solution, when the dynamic threshold setting unit constructs a dynamic threshold calculation model:
[0020] An implicit Markov model is introduced to mine the hidden state of fund transfer in network crime data, and a hidden state model of fund transfer is constructed, and then the parameters of the hidden state model of fund transfer are trained by an algorithm;
[0021] Combined with a non-parametric statistical method based on kernel density estimation, the probability density distributions of the fund transaction frequency, the involved amount, and the transaction time under different hidden states are calculated from the historical transaction data of network crimes, and the initial threshold range of the dynamic threshold calculation model is set based on this;
[0022] At the same time, the Q-learning algorithm in the reinforcement learning algorithm is used to optimize the threshold adjustment strategy of the dynamic threshold calculation model, and the accurate early warning rate is used as a reward signal to dynamically adjust the threshold.
[0023] As a further improvement of this technical solution, when the dynamic threshold setting unit dynamically adjusts the threshold according to real-time data:
[0024] A time-frequency analysis technology based on wavelet transform is used to perform multi-scale decomposition on the fund transaction flow data of network crimes to obtain the transaction frequency and the involved amount, decompose the transaction frequency and the involved amount into different frequency sub-bands, capture the short-term mutations and long-term trends in the transaction data, and design different threshold adjustment rules for different frequency sub-bands.
[0025] As a further improvement of this technical solution, when the fuzzy comprehensive evaluation unit determines the membership function:
[0026] The membership function is determined by the transaction frequency, amount concentration, and account correlation of network crime data;
[0027] Use the fuzzy C - means clustering method driven by the particle swarm optimization algorithm to initialize the transaction frequency, amount concentration, and account correlation, obtain 50 initial particles, and construct the fitness function for each particle. The initial particles represent different combinations of membership function parameters, and then iterate the fitness function of each particle to find the optimal membership function for identifying risk accounts;
[0028] At the same time, introduce the uncertainty correction mechanism of D - S evidence theory to fuse multiple groups of evidence information about the same risk account obtained from different data sources, and dynamically allocate trust weights according to the evidence conflict degree for evaluating the risk account result.
[0029] As a further improvement of this technical solution, in the weight allocation link of the fuzzy comprehensive evaluation unit:
[0030] First, use the analytic hierarchy process based on interval numbers to construct a judgment matrix, and make pairwise comparison judgments in the form of interval numbers for different cyber - crime scenarios. After consistency test and interval number operation, determine the subjective weight. Finally, based on the combined weighting model of the least - squares method, fuse the subjective weight and the objective weight, and solve the optimal combination coefficient with the minimum sum of the squares of the deviations between the two for determining the comprehensive risk score.
[0031] As a further improvement of this technical solution, when the fuzzy comprehensive evaluation unit performs fuzzy transformation to calculate the comprehensive risk score of the account:
[0032] Use the fuzzy model based on intuitionistic fuzzy reasoning to construct a rule base, and based on the intuitionistic fuzzy membership degrees of the transaction frequency, amount concentration, and account correlation, obtain the intuitionistic fuzzy evaluation result of the account risk through fuzzy implication operation and compositional reasoning.
[0033] Compared with the prior art, the beneficial effects of the present invention:
[0034] In the network security management system based on big data, the dynamic threshold setting unit constructs a dynamic threshold calculation model, adjusts the threshold according to the characteristics of cyber - crime fund flow by combining multiple algorithms, accurately adapts to the differences of different types of crimes. The fuzzy comprehensive evaluation unit uses fuzzy mathematical algorithms to accurately determine the membership function and weight, and calculate the comprehensive risk score of the account. The machine learning optimization unit uses historical data to train the model and regularly updates and optimizes it, improving the accuracy of the determination of involved accounts and enhancing the accuracy of the risk assessment of cyber - crime accounts. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is the system architecture diagram of the present invention.
[0036] The meanings of each label in the figure are as follows:
[0037] 1. Data acquisition and integration unit; 2. Dynamic threshold setting unit; 3. Fuzzy comprehensive evaluation unit; 4. Machine learning optimization unit; 5. Control unit. Detailed implementation manner
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] The present invention provides a network security management system based on big data. Please refer to Figure 1 as shown, including the following units:
[0040] The data acquisition and integration unit 1 is used to collect network crime data from multiple data sources and process the network crime data by using the DATAX multi-source heterogeneous data cleaning technology to obtain its fund flow characteristics.
[0041] When the data acquisition and integration unit 1 processes the network crime data by using the DATAX multi-source heterogeneous data cleaning technology:
[0042] There is a large amount of data with format errors, data type mismatches, and data that obviously violates business logic in multi-source heterogeneous data. The preliminary cleaning module based on pattern matching and rule constraints is used to screen out the problematic data. A predefined data format template library is established, which contains more than 50 common data formats, and more than 20 key business rules are formulated. For example, the transaction amount cannot be negative, and the user age should be within a reasonable range, etc. The multi-source heterogeneous data collected is scanned line by line, and each piece of data is matched with the data format template and business rules to screen out the data that does not meet the requirements, which reduces the burden for subsequent outlier detection and data repair. Then, the outlier detection module based on the improved isolation forest algorithm is enabled to randomly divide the feature dimensions of the remaining data and construct isolation trees. The average path length of the data point in the tree is used as the outlier score to identify the suspected outlier data. The feature dimensions of the remaining data are randomly divided into subspaces, each subspace contains some features, and 100 isolation trees with a depth of 15 are constructed. The construction process of each tree is to randomly select a feature and a splitting point to divide the data into two parts, and this process is recursively carried out until each leaf node contains only one data point or reaches the maximum depth. Calculate the average path length of each data point in all isolation trees and use it as the outlier score. The higher the outlier score, the more likely the data point is an outlier. Set an outlier score threshold to identify the data points with scores higher than the threshold as suspected outlier data, which provides a clear target for subsequent data repair. Finally, for the outlier data, the data repair module based on the generative adversarial network is used to repair the outlier data. A generative adversarial network is constructed, where the generator adopts a 5-layer convolutional neural network architecture. The transaction data over a period of time is organized into a form similar to an image matrix, with different attributes of the transaction as different channels or dimensions of the matrix, and the transaction time sequence as the rows or columns of the matrix. In this way, the convolutional neural network can perform feature extraction and processing on it like image data, making full use of the advantages of the convolutional neural network in processing multi-dimensional data. Input random noise and output repair candidate data similar to the original data distribution. The discriminator adopts a 4-layer fully connected network, which inputs the original data and the repair candidate data generated by the generator to distinguish its authenticity. The generator and the discriminator are trained adversarially. The goal of the generator is to generate repair candidate data that can deceive the discriminator, and the goal of the discriminator is to accurately distinguish the original data and the generated repair candidate data. The training process lasts for 1000 rounds. After training is completed, the generator is used to repair the suspected outlier data to obtain the repaired data. After being processed by this repair module, the repair accuracy of the data is greatly improved, providing a reliable data basis for subsequent data analysis and mining.
[0043] The acquisition process of the data acquisition and integration unit 1 has the characteristics of intelligent perception and adaptive scheduling:
[0044] In the acquisition source detection stage, a web page structure parser based on a deep neural network is used to learn and identify the data storage structures and update frequency patterns of different types of websites, construct an acquisition path planning template, collect a large number of web page samples of 5 major types of common websites, namely social platforms, e-commerce, official websites of financial institutions, news media, and government affairs websites, annotate these web pages, and the annotation content includes information such as data storage locations and data update times. A deep neural network is constructed, which includes an input layer, multiple hidden layers, and an output layer. The input layer receives the HTML code or DOM tree information of the web page. The hidden layer combines a convolutional layer and a recurrent layer. The convolutional layer is used to extract the local features of the web page, and the recurrent layer is used to process the sequential information of the web page. The output layer outputs the prediction results of the data storage structure and update frequency pattern of the website. The prepared annotated data is used to train the deep neural network, and the cross-entropy loss function and the stochastic gradient descent optimization algorithm are adopted. After multiple iterative trainings, the model can accurately learn the data storage structures and update frequency patterns of different types of websites. According to the prediction results of the trained model, more than 20 targeted acquisition path planning templates are constructed for each type of website. The templates contain information such as the starting position of data acquisition, acquisition rules, and acquisition frequencies, achieving fast and accurate positioning of the acquisition source;
[0045] When collecting from distributed data sources, a task scheduling mechanism based on the ant colony algorithm is adopted to analogize the collection tasks to the foraging paths of ants and dynamically allocate the priorities of collection tasks. The collection tasks are analogized to the foraging paths of ants, and parameters such as the number of ants, pheromone evaporation coefficient, and heuristic factor are defined. Eight key indicators such as the data update activity of the data source, network bandwidth availability, and server response latency are comprehensively considered, and a comprehensive score is calculated for each data source. Each ant selects a data source as the next collection target according to the current pheromone concentration and the comprehensive score of the data source. After completing the collection task, pheromone is released on this path. The update rule of pheromone is as follows: after pheromone evaporation, the pheromone on the path where valuable data is successfully collected increases, and the pheromone on the path where no valuable data is collected decreases. As the collection process progresses, the pheromone concentration and the comprehensive score of the data source are continuously updated, and the priorities of collection tasks are dynamically adjusted, reducing the waiting time and data redundancy in the collection process;
[0046] And it has a built-in blockchain-based trusted data caching mechanism to cope with network fluctuations and data source limitations during the collection process. The collected data fragments are stored in the form of blockchain blocks, and the collected high-value data fragments are stored in the form of blockchain blocks. Each block contains information such as data content, timestamp, and hash pointer. The hash pointer is used to link the previous block to form a chain structure of the blockchain, ensuring the immutability and traceability of the data. A buffer is set for each data source. When new data is collected, it is added to the corresponding buffer. At the same time, the buffer is regularly cleaned to delete expired or useless data. When data collection is interrupted due to network fluctuations or data source limitations, the stored data is read from the buffer to continue data collection or analysis, ensuring data continuity.
[0047] The dynamic threshold setting unit 2 constructs a dynamic threshold calculation model based on the characteristics of fund transfer, and dynamically adjusts the threshold according to the transaction frequency, involved amount, and transaction time of network crime data.
[0048] When the dynamic threshold setting unit 2 constructs the dynamic threshold calculation model:
[0049] The hidden Markov model is introduced to mine the hidden state of fund transfer in network crime data, and a hidden state model of fund transfer is constructed. Through algorithm training of model parameters, a hidden Markov model with 5 - 8 hidden states is constructed respectively. Each hidden Markov model consists of an initial state probability distribution , state transition matrix and observation probability matrix . A large amount of fund transfer data related to each type of network crime is collected. Information such as the frequency, amount, and time of fund transactions is used as the observation sequence. An algorithm is used to train the hidden Markov model. This algorithm is an iterative expectation maximization algorithm, and the specific steps are as follows:
[0050] Initialize the model parameters , and . According to the current model parameters, calculate the posterior probability of each hidden state sequence under the observation sequence, and update the model parameters , and , providing a more accurate basis for subsequent threshold setting.
[0051] Combined with a non-parametric statistical method based on kernel density estimation, estimate the probability density distributions of the frequency of fund transactions, the scale of involved amounts, and the transaction time span in different states from historical transaction data. Based on this, set an initial threshold range, extract data on the frequency of fund transactions, the scale of involved amounts, and the transaction time span under different types of cybercrimes from a large amount of historical transaction data, select a Gaussian kernel function and bandwidth parameter to perform kernel density estimation on the data of each feature (transaction frequency, scale of involved amounts, transaction time span). The formula for kernel density estimation is ; where is the estimated probability density function, is the number of data points, is the bandwidth parameter, is the kernel function, is the th data point. According to the probability density distribution obtained from kernel density estimation, determine the initial threshold range for each feature. For example, the quantiles of the distribution (such as the 90% and 95% quantiles) can be selected as the thresholds. Without relying on specific parametric distribution assumptions, it can better adapt to complex actual data and improve the rationality and accuracy of initial threshold setting;
[0052] At the same time, use the Q-learning algorithm in the reinforcement learning algorithm to optimize the threshold adjustment strategy. Take the current characteristics of fund transaction data (transaction frequency, scale of involved amounts, transaction time span) and the prediction result of the model (whether it is a suspicious transaction) as the state, define a series of threshold adjustment actions, such as increasing the threshold, decreasing the threshold, keeping the threshold unchanged, etc. Take the accurate warning rate as the reward signal, give a +10 reward for each successful identification of a potential cybercrime behavior, and give a -5 penalty for each misjudgment. Create a Q-table to store the Q-values of each state-action pair. Initially, the Q-values are all set to 0. At each time step, according to the current state select an action , and adopt -greedy strategy to randomly select an action with a probability of , and select the action with the largest Q-value with a probability of . Execute the action to obtain the next state and the reward , update the Q-values in the Q-table, repeat the above steps until the Q-table converges, and dynamically adjust the threshold according to the real-time cybercrime situation to improve the adaptability and accuracy of the system and reduce the misjudgment rate.
[0053] During the process of the dynamic threshold setting unit 2 adjusting the threshold based on real-time data:
[0054] Using time-frequency analysis technology based on wavelet transform to perform multi-scale decomposition on fund transaction flow data, decomposing the transaction frequency and the involved amount into different frequency sub-bands, capturing short-term mutations and long-term trends in the transaction data, and designing differential threshold adjustment rules for different frequency components to preprocess the real-time obtained fund transaction flow data. Select wavelet basis functions according to the characteristics of the fund transaction data. Different wavelet basis functions have different characteristics. The Daubechies wavelet basis has compact support and orthogonality, which is suitable for analyzing signals with mutation characteristics. Use the selected wavelet basis function to perform multi-scale decomposition on the preprocessed transaction frequency and involved amount data, decomposing the data into approximate components and detail components of different frequency sub-bands. For example, perform 3-layer decomposition to obtain the low-frequency approximate component High-frequency detail components , where the approximate component reflects the long-term trend of the data, and the detail component reflects the short-term changes of the data. Through the multi-scale decomposition of wavelet transform, abnormal change points in the data can be more accurately identified, providing a more accurate basis for subsequent threshold adjustment. Analyze the decomposed high-frequency detail components, and detect short-term mutations by setting a certain threshold. For example, calculate the standard deviation of each detail component. When the value of the detail component at a certain moment exceeds 3 times the standard deviation, it is determined as a short-term mutation point. Perform trend analysis on the low-frequency approximate component. Linear regression and other methods can be used to fit the change trend of the approximate component to obtain the long-term change direction and rate of the data, making the threshold adjustment more in line with the actual transaction situation and improving the early warning accuracy of the system. For the short-term mutations detected in the high-frequency detail components, if it is determined as a normal trading peak (such as increased transactions caused by promotional activities, holidays, etc.), adopt a slow change coefficient strategy. For example, the threshold is only increased by 5%-10% within 1 hour to avoid false positives due to normal short-term fluctuations. If it is determined as an abnormal fluctuation (such as abnormally large transactions, frequent small transactions, etc.), the threshold is instantly increased by 50%-80% to enhance the system's sensitivity to abnormal behaviors. For the long-term trend reflected by the low-frequency approximate component, dynamically adjust the threshold adjustment step size according to the trend slope. When the trend slope is greater than 0.3, it indicates that the transaction data shows an obvious upward trend, and the threshold is increased by 15%-25% every 6 hours. When the trend slope is less than -0.3, it indicates that the transaction data shows an obvious downward trend, and the threshold is decreased by 10%-20% every 6 hours. When the trend slope is between -0.3 and 0.3, it is considered that the transaction data is relatively stable and the threshold remains unchanged. Through differential threshold adjustment rules, the system can better cope with different types of transaction data changes and effectively reduce the false positive rate.
[0055] The fuzzy comprehensive evaluation unit 3 introduces a fuzzy mathematics algorithm to determine the membership function of network crime data, assigns weights to the transaction frequency, amount concentration, and account relevance respectively, calculates the comprehensive risk score of the network crime data account through fuzzy transformation, compares the comprehensive risk score with the set risk threshold, and determines the risk account according to the comparison result;
[0056] When the fuzzy comprehensive evaluation unit 3 determines the membership function:
[0057] Using the fuzzy C-means clustering method driven by the particle swarm optimization algorithm, 50 particles are initialized for the transaction frequency, amount concentration, and account relevance to represent different combinations of membership function parameters, and the fitness function of each particle is constructed. Then, iteration is carried out to find the optimal membership function for identifying account risks;
[0058] Collect data such as transaction frequency, amount concentration, and account relevance, and use them as input data. Initialize 50 particles, each particle representing a different combination of membership function parameters. These parameters are used to describe the cluster centers and membership matrices in fuzzy C-means clustering. Construct the fitness function of each particle, with the intra-class compactness and inter-class separation as indicators. The intra-class compactness measures the tightness between data points in the same class, and the inter-class separation measures the distance between data points in different classes. The expression of the fitness function can be designed as ; where is the weight coefficient, used to balance the importance of intra-class compactness and inter-class separation, is the intra-class compactness, is the inter-class separation, is the fitness function. The particle swarm algorithm starts to iterate. Each particle searches in the solution space according to its own velocity and position information. In each iteration, the particle updates its velocity and position. After 200 iterations, the algorithm terminates, and the combination of membership function parameters corresponding to the globally optimal position obtained at this time is the optimal membership function, which is convenient for more accurately identifying potential risk accounts;
[0059] At the same time, an uncertainty correction mechanism of D-S evidence theory is introduced to fuse multiple groups of evidence information about the same account obtained from different data sources, and trust weights are dynamically assigned according to the evidence conflict degree for evaluating the account risk result;
[0060] Obtain multiple groups of evidence information about the same account from different data sources (such as transaction records, account behavior logs, social network information, etc.), perform basic probability assignment on each group of evidence information to determine the support degree of each evidence for different risk levels (such as low risk, medium risk, high risk), and calculate the conflict degree between different evidences. The conflict degree calculation formula is ; where, is the conflict degree, and are the basic probability assignment functions of two different pieces of evidence, and are different propositions. The trust weights are dynamically assigned according to the evidence conflict degree. When the evidence conflict degree is high, the trust weights of the conflicting evidence are reduced. When the evidence conflict degree is low, the trust weights of the evidence are increased. The combination rule of D-S evidence theory is used to fuse multiple groups of evidence to obtain the final account risk assessment result, so as to effectively handle the uncertainty and conflict in multi-source evidence information, comprehensively consider the information from different data sources, and improve the credibility of the account risk assessment result.
[0061] Weighing assignment link of the fuzzy comprehensive evaluation unit 3:
[0062] First, use the analytic hierarchy process based on interval numbers to construct a judgment matrix, and make pairwise comparison judgments in the form of interval numbers for different network crime scenarios. After consistency test and interval number operation, determine the subjective weights. Make pairwise comparison judgments for each factor and give the comparison results in the form of interval numbers. For example, for factor relative to factor the importance is between and denoted as thus constructing the interval number judgment matrix where Perform a consistency test on the interval number judgment matrix to ensure the rationality of the judgment. By calculating the consistency index and the random consistency ratio when it is considered that the judgment matrix has satisfactory consistency. If not satisfied, it is necessary to re-judge and adjust the matrix. After the consistency test passes, use the interval number sorting method, such as the possibility sorting method, to determine the subjective weight interval of each factor, where is the minimum value of the subjective weight of each factor, is the maximum value of the subjective weight of each factor, and then process these intervals. The median value and other methods can be used to obtain the final subjective weight The subjective weight can better reflect the experts' judgment on the importance of each factor based on experience and knowledge, improving the reliability of weight assignment. Finally, based on the combined weighting model of the least squares method, the subjective weight and the objective weight are fused, and the optimal combination coefficient is solved with the goal of minimizing the sum of the squares of the deviations between the two, for the determination of the comprehensive risk score. The original data is standardized to eliminate the influence of the data dimension of different factors. The concept of information granularity is introduced to improve the traditional entropy weight calculation. For the th factor, its information entropy The calculation formula is adjusted after considering the information granularity, and can be specifically determined according to the actual improvement method. The objective weights of each factor are calculated according to the information entropy. , and the formula is ; where is the total number of factors, are the objective weights of each factor, is the -th factor's information entropy. The objective weights provide a data-driven basis for determining the comprehensive weights, enhancing the objectivity of weight allocation. Let the combined weight be , where is the combination coefficient, and construct the objective function ; is the objective function, is the final subjective weight, and the goal is to make minimum. Take the derivative of the objective function with respect to , and set the derivative to 0 to solve for the optimal combination coefficient . Substitute the optimal combination coefficient into the combined weight formula to obtain the comprehensive weights of each factor, which comprehensively considers subjective experience and objective data, and makes the comprehensive weights reflect the actual situation of the data by optimizing the combination coefficient.
[0063] When the fuzzy comprehensive evaluation unit 3 performs fuzzy transformation to calculate the comprehensive risk score of the account:
[0064] Construct a rule base using a fuzzy model based on intuitionistic fuzzy reasoning, taking transaction frequency, amount concentration, and account correlation as input variables, and the account risk level as the output variable. The account risk level can be divided into several levels such as low risk, medium risk, and high risk. Define corresponding intuitionistic fuzzy sets for each input and output variable. For example, for transaction frequency, intuitionistic fuzzy sets such as "low frequency", "medium frequency", and "high frequency" can be defined. Each intuitionistic fuzzy set is described by membership degree, non-membership degree, and hesitation degree. Construct a rule base containing 10 fuzzy rules based on historical data. The general form of the rule is "If (transaction frequency is A) and (amount concentration is B) and (account correlation is C), then (account risk is D)", where A, B, C, and D are intuitionistic fuzzy sets of the corresponding variables respectively. Compared with the traditional fuzzy rule base, it can capture the characteristics of account risk more accurately and provide a more comprehensive basis for subsequent risk assessment. According to the intuitionistic fuzzy membership degrees of transaction frequency, amount concentration, and account correlation, obtain the intuitionistic fuzzy evaluation result of account risk through fuzzy implication operation and compositional reasoning. According to the determined membership function, calculate the membership degree, non-membership degree, and hesitation degree of transaction frequency, amount concentration, and account correlation under their respective intuitionistic fuzzy sets. For each rule in the rule base, use the intuitionistic fuzzy implication operator for operation, perform compositional reasoning between the input intuitionistic fuzzy set and each rule in the rule base, and synthesize the reasoning results of all rules to obtain the intuitionistic fuzzy evaluation result of account risk. This result includes the membership degree, non-membership degree, and hesitation degree of account risk under each risk level, so as to more truly reflect the risk status of the account and improve the accuracy of risk assessment. Represent the intuitionistic fuzzy evaluation result of account risk in the form of an intuitionistic fuzzy set. For example, for the three levels of low risk, medium risk, and high risk, give their membership degree, non-membership degree, and hesitation degree respectively. According to the intuitionistic fuzzy evaluation result, the possibility and uncertainty of the account under different risk levels can be intuitively understood. For example, if the membership degree of the account under the high-risk level is relatively high and the hesitation degree is relatively low, it means that the account is very likely to be in a high-risk state, making the account risk assessment result more intuitive and easy to understand, helping decision-makers formulate corresponding risk response strategies according to the actual situation and improving the practicality of risk assessment.
[0065] The machine learning optimization unit 4 constructs a network crime training data set using historical case data, selects a neural network algorithm to train the network crime training data set, establishes a fund penetration threshold determination model, then regularly updates the training set with new case data, and finally determines the involved accounts in the risk accounts through the control unit 5;
[0066] Collect data related to fund penetration from channels such as historical case files and databases, covering various aspects of information such as transaction amount, transaction time, transaction object, and account association relationship, and perform cleaning. Extract features that have an important impact on fund penetration determination from the cleaned data, such as transaction frequency, concentration of fund flow, account activity, etc. According to the final determination results of historical cases, label each piece of data with whether it is involved in fund penetration, where 0 indicates not involved and 1 indicates involved. Select a neural network architecture and initialize the parameters of the neural network, including weights and biases. Divide the training dataset into a training set and a validation set, and use the training set to train the model. During the training process, calculate the output of the model through forward propagation, compare it with the true labels, calculate the error using a loss function, and then update the parameters of the model through the backpropagation algorithm to continuously optimize the performance of the model. Use the validation set to evaluate the trained model, and adjust the parameters and hyperparameters of the model according to the evaluation results until the model reaches better performance. Regularly collect new case data, integrate it with the original training dataset, and perform preprocessing operations such as cleaning, feature extraction, and label annotation on the newly added data to ensure that the format and features of the new data are consistent with the original training dataset. Use the new validation set to evaluate the retrained model, and adjust the parameters and hyperparameters of the model according to the evaluation results to ensure that the performance of the model is improved. Input the relevant data of the account to be determined into the trained fund penetration threshold determination model, including transaction records, account association information, etc. The model makes a prediction based on the input data and outputs the probability or determination result of whether the account is an involved account. The control unit 5 makes a final determination of the involved account according to the output result of the model and the preset determination rules. For example, if the involved probability output by the model exceeds 80%, the account is determined to be an involved account; if the probability is less than 20%, it is determined to be a non-involved account. For accounts with a probability between 20% and 80%, further manual review or other auxiliary means can be used for determination. Output the determination result and perform corresponding processing, such as freezing and monitoring the involved accounts, and record the determination result and relevant information for subsequent auditing and analysis, so as to be able to accurately determine involved accounts in a timely manner and effectively prevent the risk of fund penetration.
[0067] In the present invention, the multi-source data is collected and cleaned by the data acquisition and integration unit 1. The acquisition process has intelligent features and uses blockchain to cache data. The dynamic threshold setting unit 2 constructs a model to adjust the threshold for different types of cybercrimes by combining the hidden Markov model, kernel density estimation, and reinforcement learning algorithms. The fuzzy comprehensive evaluation unit 3 uses the fuzzy mathematics algorithm to determine the membership function, assign weights, and calculate the comprehensive risk score of the account. The machine learning optimization unit 4 trains and updates the model using historical data to assist in the determination of the involved accounts, effectively solving the problems of missed and misjudgments caused by fixed thresholds in traditional account risk assessments and improving the accuracy of cybercrime account risk assessments.
[0068] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A network security management system based on big data, characterized in that, It includes the following units: The data acquisition and integration unit (1) is used to collect cybercrime data from multiple data sources, and process the cybercrime data by using the DATAX multi-source heterogeneous data cleaning technology to obtain its fund flow characteristics; The dynamic threshold setting unit (2) constructs a dynamic threshold calculation model according to the fund flow characteristics, and dynamically adjusts the threshold according to the transaction frequency, involved amount and transaction time of the cybercrime data. When the dynamic threshold setting unit (2) constructs the dynamic threshold calculation model: Introduce the hidden Markov model to mine the hidden state of the fund flow in the cybercrime data, construct the hidden state model of the fund flow, and then train the parameters of the hidden state model of the fund flow through the algorithm; Combine the non-parametric statistical method based on kernel density estimation to calculate the probability density distribution of the fund transaction frequency, involved amount and transaction time under different hidden states from the historical cybercrime transaction data, and set the initial threshold range of the dynamic threshold calculation model based on this; At the same time, use the Q-learning algorithm in the reinforcement learning algorithm to optimize the threshold adjustment strategy of the dynamic threshold calculation model, and use the accurate early warning rate as the reward signal to dynamically adjust the threshold; When the dynamic threshold setting unit (2) dynamically adjusts the threshold according to the real-time data: Adopt the time-frequency analysis technology based on wavelet transform to perform multi-scale decomposition on the fund transaction flow data of cybercrime, obtain the transaction frequency and involved amount, decompose the transaction frequency and involved amount into different frequency sub-bands, capture the short-term mutations and long-term trends in the transaction data, and design different threshold adjustment rules for different frequency sub-bands; The fuzzy comprehensive evaluation unit (3) introduces the fuzzy mathematics algorithm to determine the membership function of the cybercrime data, and assigns weights to the transaction frequency, amount concentration and account relevance respectively, calculates the comprehensive risk score of the cybercrime data account through fuzzy transformation, compares the comprehensive risk score with the set risk threshold, and determines the risk account according to the comparison result; The machine learning optimization unit (4) uses the historical case data to construct a cybercrime training data set, selects the neural network algorithm to train the cybercrime training data set, establishes a fund penetration threshold determination model, then regularly updates the training set with new case data, and finally determines the involved accounts in the risk account through the control unit (5).
2. The network security management system based on big data according to claim 1, characterized in that, When the data acquisition and integration unit (1) processes the cybercrime data by using the DATAX multi-source heterogeneous data cleaning technology: Use the preliminary cleaning module based on pattern matching and rule constraints to screen out the data that does not meet the requirements, and then enable the outlier detection module based on the isolation forest algorithm to randomly divide the feature dimensions of the remaining data and construct an isolation tree, and use the average path length of the data point in the isolation tree as the outlier score to identify the outlier data; Finally, use the data repair module based on the generative adversarial network to repair the outlier data for the outlier data.
3. The network security management system based on big data according to claim 1, characterized in that, The acquisition process of the data acquisition and integration unit (1) has the characteristics of intelligent perception and adaptive scheduling: In the acquisition source detection stage, a web structure parser based on a deep neural network is used to learn and identify the data storage structure and update frequency pattern of different types of websites, and an acquisition path planning template is constructed. When collecting from distributed data sources, a task scheduling mechanism based on the ant colony algorithm is adopted to analogize the acquisition tasks to the foraging paths of ants, and the acquisition task priorities are dynamically assigned. And a trusted data caching mechanism based on blockchain is built in to cope with network fluctuations and data source limitations during the acquisition process, and the collected data fragments are stored in the form of blockchain blocks.
4. The network security management system based on big data according to claim 1, characterized in that, When the fuzzy comprehensive evaluation unit (3) determines the membership function: The membership function is determined by the transaction frequency, amount concentration, and account relevance of cybercrime data. The fuzzy C-means clustering method driven by the particle swarm optimization algorithm is used to initialize the transaction frequency, amount concentration, and account relevance, 50 initial particles are obtained and the fitness function of each particle is constructed. The initial particles represent different combinations of membership function parameters, and then the fitness function of each particle is iterated to find the optimal membership function for identifying risky accounts. At the same time, an uncertainty correction mechanism of the D-S evidence theory is introduced to fuse multiple groups of evidence information about the same risky account obtained from different data sources, and the trust weights are dynamically assigned according to the evidence conflict degree for evaluating the results of risky accounts.
5. The network security management system based on big data according to claim 1, characterized in that In the weight assignment link of the fuzzy comprehensive evaluation unit (3): First, a judgment matrix is constructed by using the analytic hierarchy process based on interval numbers, and pairwise comparison judgments are made in the form of interval numbers for different cybercrime scenarios. After consistency testing and interval number operations, the subjective weights are determined. Finally, based on the combined weighting model of the least squares method, the subjective weights and objective weights are fused, and the optimal combination coefficient is solved with the minimum sum of squared deviations between the two as the goal for determining the comprehensive risk score.
6. The network security management system based on big data according to claim 1, characterized in that When the fuzzy comprehensive evaluation unit (3) executes the fuzzy transformation to calculate the comprehensive risk score of the account: A fuzzy model based on intuitionistic fuzzy reasoning is used to construct a rule base. According to the intuitionistic fuzzy membership degrees of the transaction frequency, amount concentration, and account relevance, the intuitionistic fuzzy evaluation result of the account risk is obtained through fuzzy implication operations and compositional reasoning.
Citation Information
Patent Citations
Financial risk early warning method and system based on big data technology
CN119624661A
Systems and methods for custom ranking objectives for machine learning models applicable to fraud and credit risk assessments
US20180182029A1