Passenger ticket agent order abnormal behavior detection system
By designing a passenger ticketing agent order abnormal behavior detection system, using graph convolution network and support vector machine model, combined with deep neural network optimization correlation factors, the problem of difficulty in detecting emerging abnormal behavior in the existing technology is solved, and efficient and accurate abnormal detection effect is achieved.
Patent Information
- Application Number
- CN202510242641.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the prior art to effectively detect and identify emerging abnormal behaviors in passenger ticketing agent orders, especially ticket booking behaviors made through proxy servers bypassing IP address restrictions.
A passenger ticketing agent order abnormal behavior detection system is designed, including a data pre-acquisition module, anomaly association acquisition module, an association factor acquisition module and anomaly behavior detection model. The system constructs a detection model that can identify complex anomaly patterns through graph convolution network and support vector machine model, combined with deep neural network optimization correlation factors.
The system can effectively identify complex abnormal behavior patterns, improve the accuracy and robustness of abnormal detection, and can accurately identify abnormal orders when facing emerging fraudulent means to ensure the healthy operation of the market.
Smart Images

Figure CN120180247A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a system for detecting abnormal behaviors of passenger ticket agency orders. Background Art
[0002] With the continuous prosperity of the tourism market, the volume of passenger ticket agency business has increased significantly, which has led to an explosive growth in ticket agency order data. At the same time, it has also brought more potential risks of abnormal behaviors. For example, fraudulent ticket booking, ticket hoarding and other behaviors may disrupt the market order and damage the interests of passengers and regular ticket agencies. Therefore, an efficient abnormal behavior detection system is needed to ensure the healthy operation of the market.
[0003] In the prior art, the rules of the detection system are usually formulated based on known abnormal behavior patterns, and it is difficult to cope with continuously updated means and newly emerging abnormal behaviors. For example, in historical abnormal detection, it is stipulated that a large number of bookings for tickets on the same route within a short time using the same IP address is abnormal. However, there may be cases where this restriction is bypassed through proxy servers, and the system rules are not updated in time, resulting in such orders not being detected as abnormal, affecting the security of the system. Therefore, a system for detecting abnormal behaviors of passenger ticket agency orders is proposed here. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art and achieve the above object, the present invention proposes the following technical solution: A system for detecting abnormal behaviors of passenger ticket agency orders, comprising:
[0005] A data pre-acquisition module: collecting historical order data, dividing the order data into multiple subsets, allocating each subset to different computing nodes for hash calculation, comparing the hash values, and deleting the order data corresponding to the duplicate hash values to obtain deduplicated historical order data;
[0006] An abnormal association acquisition module: extracting the behavior characteristics of the deduplicated historical order data, collecting known abnormal order data, extracting the abnormal behavior characteristics in the abnormal order data, using the behavior characteristics and the abnormal behavior characteristics as nodes and determining the dependency relationship, constructing a graph convolutional network model, and outputting abnormal association characteristics through the model;
[0007] An association factor acquisition module: extracting association factors from the abnormal association characteristics and normal behavior characteristics, optimizing the feature representation of the association factors through a deep neural network, and constructing an abnormal detection model based on a support vector machine based on the optimized association factor feature representation;
[0008] Anomaly behavior detection model: Input the real-time order data into the anomaly detection model to obtain the anomaly results. The anomaly results are expressed as anomaly scores. Set a decision threshold. When the anomaly score is greater than the decision threshold, mark the order data corresponding to the anomaly score as abnormal and take reminder measures.
[0009] The process of collecting the historical order data is as follows:
[0010] Collect historical order data from the ticket data source. The data includes order numbers, user information, IP addresses, booking times, ticket routes, and payment information fields, and form the original order data set D = (d1, d2, d3... d n );
[0011] The process of obtaining the duplicate-removed historical order data is as follows:
[0012] Divide the original order data set D into m subsets, denoted as (D1, D2, D3... D m );
[0013] Allocate each subset to different computing nodes for hash calculation. For the order data d, calculate the hash value through an MD5 algorithm, and the calculated result is denoted as For two order data d n1 and d n2 , if then they are duplicate data. Compare the content of d n1 and d n2 again. If d n1 = d n2 , then randomly delete one of them and keep the other. After duplicate removal processing, obtain the duplicate-removed historical order data set This data set will be used as the input data for the subsequent module.
[0014] The characteristic data reflecting the order behavior includes the number of orders f1 of the user within a certain time window, the number of orders f2 under the same IP address, the number of orders f3 on the same route, the distribution characteristic f4 of the user's booking time, and the abnormal proportion f5 of the order payment method;
[0015] The process of extracting the behavior characteristics from the duplicate-removed historical order data is as follows:
[0016] Let the duplicate-removed historical order data set include i pieces of order data, denoted as Extract the characteristic data corresponding to each order data and convert the characteristic data into vector representation. The converted vector set is denoted as X = (x 1p , x 2p , x 3p ... x ip), where p is the feature dimension, p ∈ [1, 2, 3, 4, 5], and x ip represents the deduplicated historical order dataset the i-th order in the p-th feature data of. This vector set X represents the behavioral features;
[0017] The process of extracting abnormal behavioral features from abnormal order data is as follows:
[0018] Collect known abnormal order data. Let the collected abnormal order data set be A, which contains h pieces of abnormal order data, that is, A = (a1, a2, a3... a h ), extract the feature data corresponding to each abnormal order data, and convert the feature data into vector representation. The converted vector set is the abnormal order data feature set, denoted as Y = (y 1p , y 2p , y 3p ... y hp ), where y mp represents the p-th feature data of the m-th order a m in the abnormal order data set A and p ∈ [1, 2, 3, 4, 5].
[0019] The process of constructing an abnormal association model is as follows:
[0020] Combine the behavioral feature set X and the abnormal order data feature set Y into a node set, denoted as V = (x1, x2... x i , y1,, y2,... y h ), where i is the number of nodes of the behavioral features and h is the number of nodes of the abnormal order data features;
[0021] Construct the dependency relationship between nodes through a graph convolutional network. Suppose the graph convolutional network consists of L layers, and update the dependency relationship between nodes by aggregating the information of neighbor nodes;
[0022] The abnormal association model is constructed based on the graph convolutional network and is a model with an L + 1 layer network structure;
[0023] In the first layer (L = 0), the nodes are updated according to the information of direct neighbors; in the second layer (L = 1), the nodes are updated again based on the neighbor information updated in the first layer;
[0024] For the nodes x i , x i ∈ V, the feature representation in the L + 1 layer is:
[0025] where W L is the learning weight of the L-th layer, and b Lis the bias term of the L-th layer, is the feature representation of the node x in the L-th layer, c is a constant term, and σ is an activation function; i Based on the dependency relationship between nodes, a target function is constructed to train the anomaly correlation model, and the formula is expressed as:
[0026] where Q represents the true label output by the v-th node, v ∈ (1..i + h), v represents the predicted label of the v-th node;
[0027] Initialize the weights and biases in the graph convolutional network, train the model through the target function, and use the gradient descent algorithm to update the model parameters. After training and updating P times, an anomaly correlation model for predicting anomaly correlation data is obtained.
[0028] The extraction process of the correlation factor is as follows:
[0029] Extract a series of normal behavior features from the behavior features in the deduplicated historical order data
[0030] Combine the corresponding behavior features into and the anomaly correlation feature set Merge to obtain a new set C, where the dimension of C is 2j and the number of samples is N;
[0031] Suppose there are g categories in set C, which are (C1, C2... C g ), and the number of samples in each category is (e1, e2... e g ), where e g represents the number of samples in the g-th category, and where k is the index, k ∈ (1, 2,..C);
[0032] Obtain the covariance matrix of each category, and the formula is expressed as: where Σ k is the covariance matrix of the k-th category, is the sample in set C, μ k is the vector mean of category C k 1 ≤ k ≤ g, e k is the number of samples in the k-th category, and T represents the transpose;
[0033] Obtain the within-class scatter matrix S by summing the covariance matrices of all categories;
[0035] Obtain the vector means of all categories and sum them to get the total mean Based on the total mean Calculate the between-class scatter matrix, which is expressed by the formula:
[0036] Define a projection vector Z such that the result J(Z) of the formula is maximized. Obtain all Z-max that maximize J(Z). This series of projection vectors Z-max is the correlation factor, expressed as: (Z1, Z1, Z1...Z U ).
[0037] The process of optimizing the feature representation of the correlation factor through a deep neural network is as follows:
[0038] The deep neural network structure includes an input layer, a hidden layer, and an output layer;
[0039] The input data of the input layer is the correlation factor Z;
[0040] The hidden layer is expressed by the formula: where, is the weight matrix of the hidden layer, is the bias vector of the hidden layer, and σ is the activation function;
[0041] The neuron data of the output layer is U;
[0042] Train the deep neural network through a loss function and use backpropagation to calculate the gradients of the loss function with respect to the weight matrix and in the hidden layer. After P * times of propagation, obtain the optimal weight matrix and Replace the optimal weight matrix and from the hidden layer, re-enter the correlation factor Z, and each neuron in the input layer outputs the corresponding optimized correlation factor, with the number being U, expressed as
[0043] The process of constructing the anomaly detection model based on the support vector machine is as follows:
[0044] Use the optimized correlation factor as the training data, label the known normal order data and abnormal order data. The normal order data is labeled as +1, and the abnormal order data is labeled as -1 to obtain the class label corresponding to each training data
[0045] Distinguish the normal and abnormal data points in the training data through a hyperplane formula, which is expressed as: where, ωT is the transpose of the hyperplane normal vector, is the training data, and B is the intercept;
[0046] Optimize the minimum objective function through the hyperplane formula, and the optimization process is as follows:
[0047] Minimize this hyperplane formula and add a slack variable to find the optimal hyperplane to optimize the minimum objective function, which is expressed as: where η represents the penalty coefficient, U is the number of samples of the training data (the number of correlation factors), ξ is the penalty coefficient, u is the index of the training data, and u ∈ U;
[0048] At the same time, the constraint conditions need to be satisfied: ξ u > 0, where is the class label of the u-th training data;
[0049] Solve ω T and B in the objective function to obtain a decision function, which is expressed as: E * = ω T Z * + B, where Z * is the real-time input data;
[0050] Finally, an anomaly detection model based on support vector machine is obtained.
[0051] The output process of the anomaly result is as follows:
[0052] Input the new order data into the anomaly detection model in real time to calculate E * , set a decision threshold T * , when, T * > E * It is determined that the sample corresponding to the real-time input passenger ticket agency order data is abnormal. When T * ≤ E * , it is determined that the sample corresponding to the real-time input passenger ticket agency order data is normal.
[0053] The present invention has the following beneficial effects:
[0054] In the present invention, first, by collecting historical order data and performing hash calculation and duplicate removal processing, the accuracy and integrity of the data are ensured. In the case of artificially bypassing the IP address restriction through a proxy server to book tickets, the duplicate removal operation can avoid missed judgments caused by interference from duplicate data, providing a reliable data basis for accurately detecting abnormal behaviors in the subsequent process;
[0055] Secondly, by extracting various behavioral features and abnormal behavioral features and constructing the dependency relationships between nodes through a graph convolutional network, the complex associations between features can be deeply explored. When new means are artificially adopted and anomalies are simultaneously exhibited in multiple features such as booking time, price, quantity, etc., this module can more effectively identify such complex abnormal patterns by learning the dependency relationships between features, rather than being limited to the judgment of a single feature;
[0056] Finally, an SVM-based anomaly detection model is constructed through the optimized correlation factor. While fully utilizing the advantages of SVM in processing high-dimensional data and finding the optimal classification hyperplane, combined with the optimized correlation factor, the accuracy and robustness of anomaly detection are further improved. In the high-dimensional feature space, the optimized features may make the originally non-linearly separable data linearly separable, and after training with the optimized features, the correlation factor model performs more stably on new data. The combination of the two achieves a more stable effect and can effectively detect unseen abnormal patterns. In the input real-time order data, when facing newly emerging fraud means, the model can better identify abnormal orders based on the learned feature patterns and classification rules. Brief Description of the Drawings
[0057] Figure 1 It is a system block diagram of an abnormal behavior detection system for passenger ticket agency orders proposed by the present invention. Detailed Embodiments
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0059] Embodiment 1
[0060] As Figure 1 shown, an abnormal behavior detection system for passenger ticket agency orders proposed by the present invention includes:
[0061] Data pre-acquisition module: Collect historical order data, divide the order data into multiple subsets, allocate each subset to different computing nodes for hash calculation, compare the hash values, and compare and delete the order data corresponding to the duplicate hash values to obtain deduplicated historical order data;
[0062] The process of collecting historical order data is as follows:
[0063] Collect historical order data from various ticket data sources (such as ticket reservation platform databases, ticket sales system logs, etc.). These data should include but are not limited to fields such as order numbers, user information (such as names, ID numbers, contact information), IP addresses, reservation times, ticket routes, payment information, etc., to form the original order data set D = (d1, d2, d3... d n );
[0064] The process of obtaining duplicate-removed historical order data is as follows:
[0065] Divide the original order data set D into m subsets, denoted as (D1, D2, D3... D m );
[0066] Among them, by dividing the original data set into m subsets with approximately equal sizes for each subset, the data processing tasks can be evenly distributed to different computing nodes, avoiding overloading of a certain node, thereby improving the overall efficiency of data processing;
[0067] Allocate each subset to different computing nodes for hash calculation. For the order data d, calculate the hash value through an MD5 algorithm, and the calculated result is denoted as For two order data d n1 and d n2 , if then they are duplicate data. Further compare the contents of d n1 and d n2 . If d n1 = d n2 , then randomly delete one of them and keep the other;
[0068] Specifically, the steps for further comparing the contents of d n1 and d n2 are as follows:
[0069] Traverse all the contents in d n1 and d n2 . When two orders have the same other fields (user ID, reservation time, ticket route, payment amount) except for the order number, it is considered that these two orders are duplicates, and randomly delete one of them;
[0070] After duplicate removal processing, obtain the duplicate-removed historical order data set This data set will be used as the input data for the subsequent module.
[0071] Abnormal Association Acquisition Module: Extract the behavioral characteristics of deduplicated historical order data, collect known abnormal order data, extract the abnormal behavioral characteristics in the abnormal order data, use the behavioral characteristics and abnormal behavioral characteristics as nodes and determine the dependency relationship to construct an abnormal association model with a graph convolutional network structure, and output abnormal association features through the abnormal association model;
[0072] The characteristic data reflecting order behavior includes:
[0073] The number of orders f1 of the user within a certain time window, such as the number of orders in the past 1 hour, the number of orders in the past 1 day, etc., to detect abnormal high-frequency reservation behaviors;
[0074] The number of orders f2 under the same IP address, used to identify possible bulk operation behaviors;
[0075] The number of orders f3 on the same route, helping to discover abnormal reservation situations for specific routes;
[0076] The distribution characteristic f4 of the user's reservation time, such as whether it is concentrated in a certain specific time period, which may imply automated script operations;
[0077] The abnormal proportion f5 of the order payment method, for example, a large number of uses of a certain specific virtual payment method or a high proportion of attempts to pay again after unsuccessful payment, etc.;
[0078] The process of extracting the behavioral characteristics from the deduplicated historical order data is as follows:
[0079] Let the deduplicated historical order data set include i order data, denoted as Extract the characteristic data corresponding to each order data and convert the characteristic data into a vector representation. The converted vector set is denoted as X = (x 1p , x 2p , x 3p ...x ip ), where p is the feature dimension, p ∈ [1, 2, 3, 4, 5], and x ip represents the p-th characteristic data of the i-th order in the deduplicated historical order data set . This vector set X represents the behavioral characteristics;
[0080] The process of extracting the abnormal behavioral characteristics from the abnormal order data is as follows:
[0081] Collect known abnormal order data. The data sources include historical fraud cases, manually labeled abnormal data, or abnormal sample sets obtained through other reliable channels. Let the collected abnormal order data set be A, which contains h abnormal order data, that is, A = (a1, a2, a3...ah ), extract the feature data corresponding to each abnormal order data, and convert the feature data into a vector representation. The set of converted vectors is the abnormal order data feature set, denoted as Y=(y 1p , y 2p , y 3p ... y hp ), where y mp represents the p-th feature data of the m-th order a m in the abnormal order data set A, and p ∈ [1, 2, 3, 4, 5];
[0082] The process of constructing an abnormal association model is as follows:
[0083] Combine the behavior feature set X and the abnormal order data feature set Y into a node set, denoted as V=(x1, x2... x i , y1,, y 2, ... y h ), where i is the number of nodes of the behavior feature, and h is the number of nodes of the abnormal order data feature;
[0084] Construct the dependency relationship between nodes through a graph convolutional network. Suppose the graph convolutional network consists of L layers, and update the dependency relationship between nodes by aggregating the information of neighboring nodes;
[0085] The abnormal association model is constructed based on the graph convolutional network and is a model with an L+1 layer network structure;
[0086] In the first layer (L = 0), the nodes are updated according to the information of the direct neighbors; in the second layer (L = 1), the nodes are updated again based on the neighbor information updated in the first layer;
[0087] For nodes x i , x i ∈ V, the feature representation in the L+1 layer is:
[0088] where W L is the learning weight of the L-th layer, b L is the bias term of the L-th layer, is the feature representation of the node x i in the L-th layer, c is a constant term, and σ is an activation function;
[0089] Based on the dependency relationship between nodes, construct an objective function to train the abnormal association model, and the formula is expressed as: where Q v represents the true label output by the v-th node, v ∈ (1.. i+h), represents the predicted label of the v-th node;
[0090] Specifically, The calculation process of is as follows: where
[0091] represents the feature representation of the v-th node in the L-th layer of the node set; Initialize the weights and biases in the graph convolutional network, train the model through the objective function, and use the gradient descent algorithm to update the model parameters. After P times of training updates, an abnormal association model that can predict abnormal association data is obtained;
[0092] Specifically, by continuously updating the model parameters through the gradient descent algorithm, the dependence relationship between nodes can be better captured. After P iterations, the parameters of the model are optimized and can accurately reflect the dependence relationship between the behavior feature vector and the abnormal order data feature vector. At this time, the obtained model is a model that can perform abnormal association prediction. When a behavior feature vector is given, the model can output the relevant abnormal association feature vector;
[0093] Input the behavior features into the abnormal association model, and the model finally outputs a series of abnormal association features. The abnormal association feature set is represented as where represents the j-th abnormal association feature data.
[0094] Association factor acquisition module: Extract the association factors from the abnormal association features and normal behavior features, optimize the feature representation of the association factors through a deep neural network, and construct an abnormal detection model based on a support vector machine based on the optimized association factor feature representation;
[0095] The extraction process of the association factors is as follows:
[0096] Extract a series of normal behavior features from the behavior features in the deduplicated historical order data
[0097] Specifically, the normal behavior features can be the behavior features of the known normal order data in the deduplicated historical order data;
[0098] Combine the corresponding behavior feature sets into and the abnormal association feature set to obtain a new set C. Among them, the dimension of C is 2j, and the number of samples (composed of the feature data in the associated feature data set and the normal behavior feature set ) is N;
[0099] Suppose there are g categories in set C, which are (C1, C2... C g ), and the number of samples in each category is (e1, e2... eg ), where e g represents the number of samples in the g-th category, and where k is the index, k ∈ (1, 2,..C);
[0100] Obtain the covariance matrix of each category, which is expressed by the formula: where Σ k is the covariance matrix of the k-th category, is the sample in the set C, μ k is the vector mean of category C k 1 ≤ k ≤ g, e k is the number of samples in the k-th category, T represents transpose;
[0101] Obtain the within-class scatter matrix S by summing the covariance matrices of all categories;
[0102] Obtain the vector means of all categories and sum them to obtain the total mean Based on the total mean Calculate the between-class scatter matrix, which is expressed by the formula:
[0103] Define a projection vector Z such that the result J(Z) of the formula is maximized. Obtain all Z that maximize J(Z). This series of projection vectors Z is the correlation factor, denoted as (Z1, Z1, Z1...Z U );
[0104] Specifically, to find Z that maximizes J(Z), this process is equivalent to solving the generalized eigenvalue problem: where λ is the generalized eigenvalue. By performing eigenvalue decomposition on the matrix, a series of projection vectors (Z1, Z1, Z1...Z U ) are obtained. These eigenvectors are sorted according to the magnitude of the eigenvalues, U ≤ 2j, where 2j is the dimension of the set C;
[0105] The process of optimizing the feature representation of the correlation factor through a deep neural network is as follows:
[0106] The deep neural network structure includes an input layer, a hidden layer, and an output layer;
[0107] The input data of the input layer is the correlation factor Z;
[0108] The hidden layer is expressed by the formula: where is the weight matrix of the hidden layer, is the bias vector of the hidden layer, and σ is the activation function;
[0109] The data of the output layer neurons is U (the types of output data, one data represents one neuron);
[0110] Train the deep neural network through a loss function and use backpropagation to calculate the gradients of the loss function with respect to the weight matrices and in the hidden layer. After P * propagations, obtain the optimal weight matrices and Replace the optimal weight matrices and in the hidden layer, re-enter the correlation factor Z, and each neuron in the input layer outputs the corresponding optimized correlation factor, with the number being U, denoted as
[0111] The construction process of the anomaly detection model based on the support vector machine is as follows:
[0112] Use the optimized correlation factor as the training data, label the known normal order data and abnormal order data, label the normal order data as +1, and label the abnormal order data as -1 to obtain the class label corresponding to each training data
[0113] Distinguish the normal and abnormal data points in the training data through a hyperplane formula, which is expressed as: where ω T is the transpose of the hyperplane normal vector, is the training data, and B is the intercept;
[0114] Optimize the minimum objective function through the hyperplane formula, and the optimization process is as follows:
[0115] Minimize this hyperplane formula and add a slack variable to find the optimal hyperplane to optimize the minimum objective function, which is expressed as: where η represents the penalty coefficient, U is the number of samples of the training data (the number of correlation factors), ξ is the penalty coefficient, u is the index of the training data, and u ∈ U;
[0116] At the same time, it is necessary to satisfy the constraint conditions: ξ u > 0, where is the class label of the u-th training data;
[0117] After optimizing the minimum objective function through the hyperplane formula, solve ω T and B in the objective function. Directly use the Lagrange duality principle to transform the original problem into a dual problem for solution, and obtain a decision function, which is expressed as: E * = ωT Z * + B, where Z * is real-time input data, ω T and B in the formula are values obtained after solving, satisfying the above constraints ξ u > 0. This decision function is the formulaic representation of the anomaly detection model, which can receive new sample data (real-time order data) and assign an anomaly score (the output value of the decision function, i.e., the anomaly score). In this way, an anomaly detection model for distinguishing abnormal data constructed based on the optimized correlation factor is obtained.
[0118] Anomaly behavior detection model: Input real-time order data into the anomaly detection model to obtain anomaly results. The manifestation of the anomaly results is the anomaly score. Set a decision threshold. When the anomaly score is greater than the decision threshold, mark the order data corresponding to the anomaly score as abnormal and take reminder measures;
[0119] The output process of the anomaly results is as follows:
[0120] When new order data is input in real-time, input it into the anomaly detection model, that is, substitute it into the decision function E * = ω T Z * + B to calculate E * , and set a decision threshold T * . If T * > E * , then it is determined that the sample corresponding to the real-time input passenger ticketing agency order data is abnormal because it is less than the decision threshold T * . In the model, it indicates that the order is on the "abnormal side" of the hyperplane defined by the model. For example, there may be situations such as a sudden large increase in the number of tickets booked, an abnormal deviation of the ticket price from the normal range, or an abnormal booking time that does not conform to the normal pattern;
[0121] If T * ≤ E * , then it is determined that the sample corresponding to the real-time input passenger ticketing agency order data is normal, which means that the order is on the "normal side" of the hyperplane and conforms to the characteristics of a normal order. For example, the number of tickets booked, the price, and the time are all within the normal fluctuation range;
[0122] Among them, the decision threshold T * is obtained by screening out the decision function values E * corresponding to normal orders (marked as 0) among the calculated decision function values, and calculating the mean and standard deviation of the decision function values of normal orders. For example, assume that the decision function values of normal orders are [0.2, 0.3, 0.1, 0.4,...], and the calculated mean is 0.25 and the standard deviation is 0.1;
[0123] According to business requirements and the tolerance for false alarm rates, select a constant (such as 2 or 3), set the decision threshold as the mean minus the constant times the standard deviation, and the final obtained value is the decision threshold T * ;
[0124] After being judged as abnormal, the system triggers a security alarm to prompt the network administrator of the abnormal order behavior, and the administrator takes further investigation or preventive measures.
[0125] Specifically, the model distinguishes normal and abnormal data points by finding a hyperplane, rather than based on fixed rules formulated by humans. For example, when faced with a new means artificially adopted to bypass existing rules, the data-driven method on which the model is based can automatically capture new abnormal patterns. Suppose a new strategy is artificially adopted to conduct abnormal ticket booking through complex ticket booking operations while avoiding IP address detection. Since the model is trained based on the correlation between a large amount of historical order data (including normal and abnormal orders) and optimizes this correlation by combining a deep neural network, it can better learn the distribution of order data in the high-dimensional feature space. Therefore, when a new abnormal behavior appears, even if this behavior does not conform to the old rules, it will be divided into the abnormal area by the hyperplane.
[0126] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A passenger ticketing agency order abnormal behavior detection system, characterized in that: include: Data pre-acquisition module: collects historical order data, divides the order data into multiple subsets, assigns each subset to different computing nodes for hash calculation, compares hash values, deletes order data corresponding to duplicate hash values, and obtains deduplicated historical order data; Abnormal correlation acquisition module: extracts the behavioral features of deduplicated historical order data, collects known abnormal order data, extracts abnormal behavior features from abnormal order data, uses behavioral features and abnormal behavior features as nodes and determines dependencies, builds a graph convolutional network model, and outputs abnormal correlation features; Correlation factor acquisition module: extracts correlation factors from abnormal correlation features and normal behavior features, optimizes the feature representation of correlation factors through deep neural networks, and builds an anomaly detection model based on support vector machines based on the optimized correlation factor feature representation; Abnormal behavior detection model: Real-time order data is input into the anomaly detection model to obtain abnormal results. The abnormal results are expressed as an anomaly score. A decision threshold is set. When the anomaly score is greater than the decision threshold, the order data corresponding to the anomaly score is marked as abnormal, and reminder measures are taken.
2. A passenger ticket agency order abnormal behavior detection system according to claim 1, characterized in that: The historical order data collection process is as follows: Collect historical order data from the ticketing data source, including order number, user information, IP address, booking time, ticket route, and payment information fields, and form the original order data set D = (d1, d2, d3...d n ); The process of obtaining deduplicated historical order data is as follows: The original order data set D is divided into m subsets, represented as (D1, D2, D3...D m ); Each subset is assigned to a different computing node for hash calculation. For the order data d, the hash value is calculated using an MD5 algorithm. The result after calculation is expressed as For two order data d n1 and d n2 ,when When , it is expressed as repeated data, and d is compared again n1 and d n2 If the content of n1 =d n2 , then randomly delete one of them and keep the other one. After deduplication processing, we get the deduplication historical order data set and use it as input data.
3. A passenger ticket agency order abnormal behavior detection system according to claim 2, characterized in that: The process of extracting behavioral features from deduplicated historical order data is as follows: Set up a deduplicated historical order data set It includes i pieces of order data, expressed as Extract the feature data corresponding to each order data, and convert the feature data into vector representation. The converted vector set is represented as X = (x 1p , x 2p , x 3p ...x ip ), where p is the feature dimension p∈[1, 2, 3, 4, 5], x ip Represents a dataset of deduplicated historical orders The i-th order The pth feature data, the vector set X represents the behavior feature set; The process of extracting abnormal behavior features from abnormal order data is as follows: Collect known abnormal order data. Suppose the collected abnormal order data set is A, which contains h abnormal order data, that is, A = (a1, a2, a3...a h ), extract the feature data corresponding to each abnormal order data, and convert the feature data into a vector representation. The converted vector set is the abnormal order data feature set, which is expressed as Y = (y 1p ,y 2p ,y 3p ...y hp ), where y mp Represents the mth order a in the abnormal order dataset A m The pth feature data and p∈[1, 2, 3, 4, 5].
4. A passenger ticket agency order abnormal behavior detection system according to claim 3, characterized in that: The process of constructing the abnormal association model is as follows: The behavior feature set X and the abnormal order data feature set Y form a node set, expressed as V = (x1, x2...x i ,y1,,y2,...y h ), where i is the number of nodes of behavioral features, and h is the number of nodes of abnormal order data features; The dependency relationship between nodes is constructed through the graph convolution network. Assume that the graph convolution network structure includes L layers, and the dependency relationship between nodes is updated by aggregating the information of neighboring nodes; The abnormal association model is built based on the graph convolutional network, which is a model with an L+1 layer network structure; In the first layer, L = 0, nodes are updated based on the information of their direct neighbors; in the second layer, L = 1, nodes are updated again based on the updated neighbor information in the first layer; For node x i , x i ∈V, the feature representation at the L+1 layer is: Among them, W L Learn weights for layer L, b L is the L-th layer bias term, is the L-th layer node x i The characteristic representation of is, c is a constant term, and σ is an activation function; Based on the dependency relationship between nodes, an objective function is constructed to train the abnormal association model. The formula is expressed as: Among them, Q v represents the true label output by the vth node, v∈(1, 2, 3....i+h), Represents the predicted label of the vth node; Initialize the weights and bias items in the graph convolutional network, train the model through the objective function, and use the gradient descent algorithm to update the model parameters. After training and updating P times, an abnormal correlation model for predicting abnormal correlation data is obtained.
5. A passenger ticket agency order abnormal behavior detection system according to claim 3, characterized in that: The extraction process of the correlation factor is: Extract normal behavior features from behavior features in deduplicated historical order data The corresponding behavioral features are grouped as and abnormal associated feature set Merge to obtain a new set C, where the dimension of C is 2j and the number of samples is N; Suppose there are g categories in set C, namely (C1, C2...C g ), the number of samples in each category is (e1, e2...e g ), where e g represents the number of samples in the g-th category, and Where k is the index, k∈(1, 2, ..C); Get the covariance matrix for each category, the formula is expressed as: Among them, Σ k is the covariance matrix of the kth category, is a sample in set C, μ k Category C k The vector mean of 1≤k≤g, e k is the number of samples in the kth category, and T represents transposition; The intra-class scatter matrix S is obtained by summing the covariance matrices of all classes; Get the vector mean of all categories and sum them to get the sum of the means Based on the sum of the mean Calculate the inter-class scatter matrix, the formula is expressed as: Define a projection vector Z such that the formula The result J(Z) is the largest, and all Z that make J(Z) the largest are obtained. The largest projection vector Z is the correlation factor, which is expressed as: (Z1, Z1, Z1...Z U ).
6. A passenger ticket agency order abnormal behavior detection system according to claim 1, characterized in that: The process of optimizing the feature representation of the correlation factor through a deep neural network is as follows: The deep neural network structure contains input layer, hidden layer and output layer; The input data of the input layer is the correlation factor Z; The hidden layer formula is expressed as: in, is the weight matrix of the hidden layer, is the bias vector of the hidden layer, σ is the activation function; The output layer neuron data is U; A deep neural network is trained by a loss function and back-propagation is used to calculate the loss function for the weight matrix in the hidden layer and The gradient of * After propagation, the optimal weight matrix is obtained and The optimal weight matrix and Replace from the hidden layer, re-input the correlation factor Z, and each neuron in the input layer outputs the corresponding optimized correlation factor, the number is U, expressed as 7. A passenger ticket agency order abnormal behavior detection system according to claim 1, characterized in that: The construction process of the anomaly detection model based on support vector machine is as follows: The optimized correlation factor As training data, mark the known normal order data and abnormal order data. Normal order data is marked as +1, and abnormal order data is marked as -1. Obtain the category label corresponding to each training data A hyperplane formula is used to distinguish normal and abnormal data points in the training data. The formula is expressed as: Among them, ω T is the transpose of the hyperplane normal vector, is the training data, B is the intercept; The minimum objective function is optimized by the hyperplane formula, and the optimization process is: Minimize this hyperplane formula and add a slack variable to find the optimal hyperplane to optimize the minimum objective function, the formula is expressed as: Among them, η represents the penalty coefficient, U is the number of sample size correlation factors of the training data, ξ is the penalty coefficient, u is the index of the training data, u∈U; At the same time, the constraints must be met: ξ u >0, where is the category label of the u-th training data; Solving the objective function ω T and B to obtain a decision function, the formula is expressed as: E * =ω * Z - +B, where Z * To input data in real time; Finally, an anomaly detection model based on support vector machine is obtained.
8. A passenger ticket agency order abnormal behavior detection system according to claim 7, characterized in that: The output process of the abnormal result is: Input new order data into the anomaly detection model in real time to calculate E * , set a decision threshold T * , when T * >E * , it is judged that the sample corresponding to the real-time input passenger ticketing agent order data is abnormal. When T * ≤E * , it is determined that the sample corresponding to the real-time input passenger ticketing agency order data is normal.
Citation Information
Cited By
Bill collaborative management method and system based on multi-source heterogeneous data fusion
CN120450882A
E-commerce platform order analysis system and method based on data mining
CN121616379A