Ticket order anomaly detection system based on machine learning
By adopting machine learning-based methods in the ticket order exception detection system, including feature extraction, dynamic threshold adjustment and cross-domain detection, the problems of feature neglect, large calculation overhead and non-dynamic threshold in the existing system are solved, and more efficient and flexible abnormal detection effects are achieved.
Patent Information
- Application Number
- CN202510044289.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-11
- Publication Date
- 2025-05-30
AI Technical Summary
The existing ticket order abnormality detection system has problems such as feature extraction methods ignore rare but important features, over-reliance on common high-frequency features, large calculation overhead, undynamic threshold settings, inability to comprehensively evaluate user behavior, lack of cross-domain detection and inflexible optimization processes.
It adopts a ticket order abnormality detection system based on machine learning, including data collection, data preprocessing, exception rule engine, exception order detection and exception early warning module. Feature extraction and processing are performed through density clustering algorithm, TF-IDF and Laplace smoothing, adaptive thresholds and blacklist scores are dynamically adjusted, cross-domain rule detection is implemented, and hyperparameters are optimized through cross-entropy loss function and love force formula.
It increases the attention to rare but important features, reduces calculation overhead, enhances the flexibility and accuracy of detection, effectively reduces false positives and missed reports, and promptly identify potential fraud.
Smart Images

Figure CN120069165A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of order detection, and more specifically, to a ticket order anomaly detection system based on machine learning. Background Art
[0002] A patent with the publication number CN116934418A discloses a method, system, device, and storage medium for detecting and warning abnormal orders. The method includes screening and eliminating the order lists of each promotion period at the big promotion node to obtain the order transaction data of each promotion period; according to the key information and the order transaction data, through the calculation of the order correlation degree, extracting the key order data of each promotion period, and summarizing to obtain the key order data of the big promotion node; inputting the key order data of the big promotion node into the order anomaly detection model to obtain the order anomaly detection result; according to the order anomaly detection result, extracting all abnormal order users, and screening all abnormal order users according to the commodity purchase data of all abnormal order users to obtain the final abnormal order users; and performing abnormal order warning according to the key order data of the big promotion node and the final abnormal order users. This embodiment realizes the rapid detection and warning of abnormal orders under the big promotion node, and improves the accuracy of abnormal order detection.
[0003] Existing ticket order anomaly detection systems mainly have the following problems:
[0004] Only relying on traditional feature extraction methods may ignore some rare but important features. The model may overly rely on frequently occurring features and ignore those key but scarce features, which may affect the prediction accuracy and generalization ability of the model; without dynamically optimizing TF-IDF and performing smoothing adjustment, it may overly rely on some common high-frequency features, lack attention to low-frequency important features, and cause the model to overfit the features in the training set; without a smoothing mechanism and adaptive parameter adjustment, the model may process a large amount of useless or noisy data during calculation, resulting in unnecessary computational overhead and affecting efficiency and the response speed of the model;
[0005] The model only relies on a fixed threshold for anomaly detection and may not be able to adapt to the behavior patterns of different users. For some users, the orders are frequent and the amounts are large, while for other users, the order frequency is low and the amounts are small. The fixed threshold may lead to misjudgment or missed judgment of normal transactions. Lack of a mechanism to dynamically adjust the threshold based on the user's historical behavior easily results in the inability to effectively identify potential abnormal transactions, increasing the risk of false positives or false negatives. Without considering the blacklist score, the system may not be able to comprehensively evaluate multiple dimensions of user behavior, thus unable to accurately identify users who frequently conduct abnormal transactions. Lack of a comprehensive scoring mechanism may cause the system to fail to timely identify users with fraud risks, resulting in potential fraud behaviors not being effectively detected. Without cross-domain rules, the system may not be able to timely detect the behavior of a user making payments using the same account on different devices or IP addresses. The lack of detection of these cross-domain behaviors may lead to the system's inability to timely identify fraud activities, increasing the risk of system abuse.
[0006] Without a cross-entropy loss function to measure individual fitness, the algorithm may not be able to accurately evaluate the pros and cons of each hyperparameter configuration, resulting in an inefficient optimization process. It may require more computing resources and time to find the optimal hyperparameters. In addition, the lack of dynamic adjustment of the historical trend factor makes the optimization process unable to effectively capture the change trend of individual performance, thereby affecting the optimization efficiency and leading to stagnation or slow convergence speed. Without using the love force formula to measure the intimacy between individuals, the optimization process may be too random or mechanical, resulting in insufficient exploration of the search space between individuals and potentially missing excellent hyperparameter combinations. The lack of reasonable intimacy update will lead to insufficient cooperation between individuals in the optimization process, reducing the global optimization ability and unable to fully utilize existing good solutions at the right time. Without considering the dynamic adjustment of the weight of the trend factor, the algorithm may overly rely on early performance during the optimization process, resulting in unreasonable weight settings and thus affecting the fitness evaluation. This leads to a lack of refinement and flexibility in the optimization process, making the algorithm unable to efficiently explore the optimal solution when facing complex data. Without considering controlling the adjustment rate of the trend factor, it may cause the performance of the optimization process to be inconsistent at different stages and unable to dynamically adapt to changes in task and data characteristics.
[0007] In view of this, the present invention proposes a ticket order anomaly detection system based on machine learning to solve the above problems. Summary of the Invention
[0008] To overcome the above defects of the prior art and to achieve the above object, the present invention provides the following technical solution: A ticket order anomaly detection system based on machine learning, including:
[0009] A data collection module for collecting ticket-related data;
[0010] A data preprocessing module for preliminarily processing the collected ticket-related data to obtain a comprehensive ticket feature dataset;
[0011] An anomaly rule engine module for performing anomaly detection on the comprehensive ticket feature dataset according to preset rules to obtain a suspected anomaly order feature dataset;
[0012] An anomaly order detection module for training to obtain an anomaly order prediction model based on the suspected anomaly order feature dataset, predicting an anomaly order evaluation label based on the anomaly order prediction model, and determining whether a ticket order is abnormal;
[0013] An anomaly warning module. If a ticket order is abnormal, the order anomaly detection terminal generates a warning message and sends a warning notice to relevant personnel; each module is connected by wired and / or wireless means.
[0014] Furthermore, the ticket-related data includes user information data, order information data, payment information data, and user behavior data;
[0015] The user information data includes the user login IP, geographical location, and account information; the order information data includes the ticket purchase time, ticket type, ticket price, number of tickets purchased, and performance; the payment information data includes the payment method, payment amount, and payment time; the user behavior data includes the ticket purchase frequency, refund times, and operation duration.
[0016] Furthermore, the method for preliminarily processing the collected ticket-related data to obtain a comprehensive ticket feature dataset includes:
[0017] Identifying and removing outliers in the user information data, order information data, payment information data, and user behavior data through a density clustering algorithm to obtain a user information feature dataset, an order information feature dataset, a payment information feature dataset, and a user behavior feature dataset;
[0018] Performing feature extraction on the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset through TF-IDF and Laplace smoothing, and performing standard deviation normalization processing on the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset after feature extraction, converting them into a standard normal distribution with a mean of 0 and a standard deviation of 1 to obtain a normalized user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset;
[0019] Fusing the normalized user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset through a weighted model to obtain a comprehensive ticket feature dataset.
[0020] Further, the method for feature extraction of the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset through TF-IDF and Laplace smoothing includes:
[0021] For the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset, calculate their term frequencies and inverse document frequencies respectively through the term frequency calculation formula and the inverse document frequency calculation formula;
[0022] The term frequency calculation formula is: where TF(ω j , T i ) is the number of occurrences of the feature ω j in the document T i ; ω j is a certain feature in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset; T i is a certain document in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset; Co(ω j , T i ) is the number of occurrences of the feature ω j in the document T i ; N is the total number of occurrences of all features in the document; j is the index of the feature; i is the index of the document;
[0023] The inverse document frequency calculation formula is: where D is the set of all documents; m is the total number of documents in the document set; n j is the number of documents containing the feature;
[0024] Calculate the TF-IDF value of each feature, and generate a feature vector for each document. The TF-IDF value calculation formula is: TF-IDF(ω j , T i , D) = TF(ω j , T i ) · IDF(ω j , D); where TF-IDF(ω j , T i , D) is the TF-IDF value of the feature ω j in the document T i ;
[0025] Smooth the term frequency of each feature through the Laplace smoothing formula. The Laplace smoothing formula is: Among them, TF sm (ω j , T i ) is the word frequency value after Laplace smoothing; α d is the adaptive smoothing parameter; V is the total number of all features in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset;
[0026] The adaptive smoothing parameter is dynamically designed through the adaptive smoothing parameter formula, and the adaptive smoothing parameter formula is: Among them, c 1 is the influence factor of the total number of occurrences N of all features in the document on the smoothing parameter; c 2 is the influence factor of the total number of all features V in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset on the smoothing parameter; c 3 is the constant factor for adjusting the overall smoothing effect.
[0027] Adjust the word frequency of each feature through the above calculation and the Laplace smoothing formula, and finally combine the TF-IDF values of the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset to generate a multi-dimensional feature vector.
[0028] Furthermore, the method for fusing the normalized user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset through a weighted model to obtain a comprehensive ticket feature dataset includes:
[0029] Denote the user information feature dataset as F 1 , the order information feature dataset as F 2 , the payment information feature dataset as F 3 , and the user behavior feature dataset as F 4 ; the weighted model is: EPR = F 1 ·δ 1 + F 2 ·δ 2 + F 3 ·δ 3 + F 4 ·δ 4 ; where EPR is the comprehensive ticket feature dataset; δ 1 is the weight coefficient of the user information feature dataset; δ 2 is the weight coefficient of the order information feature dataset; δ 3 is the weight coefficient of the payment information feature dataset; δ 4 is the weight coefficient of the user behavior feature dataset.
[0030] Further, the preset rules include a threshold rule, a blacklist matching rule, and a cross-domain rule.
[0031] Further, the method for performing anomaly detection on the comprehensive ticket feature dataset according to the preset rules to obtain a suspected anomaly order feature dataset includes:
[0032] Performing anomaly detection on the comprehensive ticket feature dataset according to the threshold rule through an adaptive threshold formula, and the adaptive threshold formula is: T u = β 1 · log(N u + 1)+ β 2 · avg(O u ); where, T u is the adaptive threshold of user u; N u is the number of orders of the user within the past T win time window; O u is the order amount of user u within the past T win time window; avg(O u ) is the average value of the order amounts of user u within the past T win time window; β 1 is the purchase behavior adjustment coefficient; β 2 is the average order amount adjustment coefficient; u is the index of the user;
[0033] According to the threshold rule, if the order amount or the number of orders of a certain user exceeds the adaptive threshold T u of this user, then it is determined that this order is a suspected abnormal order;
[0034] Calculating the behavior score of the blacklist user through the blacklist scoring calculation formula, and the blacklist scoring calculation formula is: where, S b is the blacklist score of user u; is the purchase frequency of user u; is the order amount volatility of user u; σ(O u ) is the standard deviation of the order amount of user u; is the order speed; T ta is the average order interval time of user u; γ 1 is the purchase frequency adjustment coefficient; γ 2 is the order amount volatility adjustment coefficient; γ 3 is the order speed adjustment coefficient;
[0035] According to the blacklist matching rule, preset a blacklist score threshold. When the blacklist score of user u exceeds the blacklist score threshold, then it is determined that this user is a blacklist user;
[0036] According to the cross - domain rules, if user u switches different IP addresses or different devices within the past T win time window and pays for N u orders using the same account, it is determined that the user has abnormal payment behavior.
[0037] Furthermore, the training method of the abnormal order prediction model includes:
[0038] Divide the data set into a training set, a validation set, and a test set, and construct an abnormal order prediction model; the sample set is a subset of the data set, and each sample set includes a historical suspected abnormal order feature data set and a corresponding abnormal order evaluation label;
[0039] The input data of the model is the historical suspected abnormal order feature data set, the output label of the model is the abnormal order evaluation label, and the abnormal order evaluation label is divided into two categories: abnormal orders and normal orders; use the sigmoid function as the activation function; the abnormal order prediction model is a gradient - boosting decision tree model;
[0040] Initialize the abnormal order prediction model, set the hyperparameters of the number of trees, learning rate, and tree depth; through k - fold cross - validation, use different combinations of hyperparameters to train the model, and use the results of cross - validation to tune the initially set hyperparameters, and select the best - performing combination of hyperparameters; use binary cross - entropy as the loss function to measure the error of model prediction;
[0041] Train the abnormal order prediction model on the training set, and train a new decision tree based on the residuals of the current model in each iteration; according to the model performance feedback, adjust the hyperparameters of the model to optimize the model, retrain the model with the adjusted hyperparameters, and after the training process ends, select the optimal hyperparameters through cross - validation;
[0042] Stop when the training reaches the preset model complexity or convergence condition, and obtain the finally trained abnormal order prediction model; use the trained abnormal order prediction model to predict the current suspected abnormal order feature data set to obtain the abnormal order evaluation label.
[0043] Furthermore, the method for adjusting the hyperparameters of the model to optimize the model includes:
[0044] S91. Randomly initialize the population, generate E individuals, and form the initial population G = {x 1 , x 2 ,..., x p ,..., x E}; where x p is the p - th individual in the population; E is the total number of individuals; p is the index of the individual, p ∈ {1,..., E};
[0045] S92. There are d hyperparameters in the preset model, and each individual in the population represents a set of hyperparameter configurations; each individual is represented as a d-dimensional vector: x p =(x p1 ,x p2 ,...,x pd ); where x pd is the value of the d-th hyperparameter of the p-th individual;
[0046] S93. The fitness of an individual is measured by the binary cross-entropy loss function, and the current fitness f(x p ) of each individual is evaluated; the historical trend factor of each individual is calculated: Δf(x p ) = f(x p ) - f -1 (x p ); where Δf(x p ) is the historical trend factor of each individual x p , representing the fitness improvement value of the individual x p in this round of iteration; f -1 (x p ) is the fitness of the individual x p in the previous round of iteration;
[0047] S94. Based on the fitness and the distance between individuals, the love force is calculated through the love force formula to measure the intimacy between two individuals; the love force formula is: where L(x p ,x q ) is the love force between the individual x p and the individual x q ; ds(x p ,x q ) is the distance between the individual x p and the individual x q ; f(x q ) is the fitness of the individual x q ; Δf(x q ) is the historical trend factor of the individual x q ; λ is the trend factor weight; the trend factor weight λ is dynamically adjusted through the trend factor weight adjustment formula;
[0048] S95. Each individual updates its intimacy through the intimacy update formula according to the love force with other individuals; the intimacy update formula is: where Q(x p ,t + 1) is the intimacy of the individual x p at the (t + 1)-th iteration; Q(x p, the intimacy of individual x at the current t-th iteration; p At the current t-th iteration;
[0049] S96. A preset intimacy threshold. When the intimacy between two individuals x p and x q reaches the preset intimacy threshold, mating is performed through the mating formula to generate a new individual; the mating formula is: x new = a·x p +(1 - a)·x q ; where x new is the new individual; a is a random factor that controls the mating degree of the two individuals and ranges between 0 and 1;
[0050] S97. Perform fitness evaluation on the newly generated individual. If the fitness of the new individual is better, replace the original individual; otherwise, keep the original individual unchanged; a preset fitness threshold. Repeat S94 - S97 until the maximum number of iterations is reached or the preset fitness threshold is reached, and then stop. Finally, output the optimal hyperparameter configuration of the model, that is, the individual with the optimal fitness.
[0051] Furthermore, the trend factor weight adjustment formula is: where λ′ is the adjusted trend factor weight; t is the current iteration number; t max is the maximum number of iterations; b is an exponential factor that controls the adjustment rate of the trend factor weight.
[0052] Technical effects and advantages of the ticket order anomaly detection system based on machine learning of the present invention:
[0053] The present invention can quantify the importance of each feature through the combination of TF-IDF; TF (term frequency) reflects the frequency of a feature in a single document, and IDF (inverse document frequency) suppresses the influence of common features by reducing the weight of features that frequently appear in multiple documents, thereby enhancing the attention to rare but important features; Laplace smoothing solves the situation where the term frequency is zero by adjusting the frequency of features, avoiding calculation problems in the model caused by some features not appearing in the document, and can smooth the feature values, enabling the model to still maintain effective learning when facing rare features; by dynamically adjusting the adaptive smoothing parameter, the feature smoothing effect can be dynamically optimized according to the density of the document content or the distribution of features, improving the flexibility of the model;
[0054] The adaptive threshold formula can dynamically calculate the threshold for each user based on the number of orders and the amount within a past time window, so as to adjust the detection sensitivity according to the historical behavior habits of different users. By setting the threshold in this dynamic, user behavior-driven way, abnormal orders can be detected more accurately, avoiding misjudgments caused by fixed threshold settings; by dynamically adjusting the threshold and multi-dimensional anomaly detection rules, the false positive rate and false negative rate can be effectively reduced; the blacklist score can comprehensively evaluate the behavior of users. Combining various factors such as the user's purchase frequency, order amount volatility, and order speed, it can accurately identify those users who frequently conduct abnormal transactions. By calculating the blacklist score, the system can comprehensively analyze the behavior of users according to the score level to determine whether there is potential fraud behavior; the cross-domain rule can effectively detect the situation where a user uses the same account to make payments on different IP addresses or devices, which is a typical abnormal payment behavior. Through this cross-domain rule, risks such as identity theft and payment fraud by users can be discovered in a timely manner;
[0055] By randomly initializing the population and using the cross-entropy loss function to measure the fitness of individuals, different hyperparameter configurations can be efficiently explored to find the optimal hyperparameters suitable for the current task; this search method, by simulating the process of natural selection, can not only optimize a single hyperparameter but also adjust multiple hyperparameters simultaneously, thus greatly improving the overall performance of the model; through the historical trend factor and the love force formula, the optimization process of individuals and the intimacy between individuals can be dynamically evaluated, ensuring that during the optimization process, the intimacy between individuals is reasonably measured and adjusted, enabling the model to flexibly adapt to data changes during the iteration process; the dynamic adjustment of the trend factor weight makes the optimization process more meticulous. As the optimization process progresses, the weight factor can more reasonably affect the fitness evaluation of individuals, enhancing the accuracy of the search; by introducing an exponential factor to control the adjustment rate of the trend factor, the optimization process of the algorithm can dynamically adapt according to different task and data characteristics. Brief Description of the Drawings
[0056] Figure 1 It is a schematic structural diagram of the ticket order anomaly detection system based on machine learning of the present invention;
[0057] Figure 2 It is a schematic flow diagram of the ticket order anomaly detection method based on machine learning of the present invention. Detailed Embodiments
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] Embodiment 1
[0060] Please refer to Figure 1 As shown in the figure, the ticket order anomaly detection system based on machine learning in this embodiment includes:
[0061] A data acquisition module, which is used to acquire ticket-related data;
[0062] A data preprocessing module, which is used to perform preliminary processing on the acquired ticket-related data to obtain a comprehensive ticket feature data set;
[0063] An anomaly rule engine module, which is used to perform anomaly detection on the comprehensive ticket feature data set according to preset rules to obtain a suspected anomaly order feature data set;
[0064] An anomaly order detection module, which is used to train and obtain an anomaly order prediction model based on the suspected anomaly order feature data set, predict an anomaly order evaluation label based on the anomaly order prediction model, and determine whether the ticket order is abnormal;
[0065] An anomaly warning module. If the ticket order is abnormal, the order anomaly detection terminal generates a warning message and sends a warning notice to relevant personnel; each module is connected by wired and / or wireless means.
[0066] The ticket-related data includes user information data, order information data, payment information data, and user behavior data;
[0067] The user information data includes the user login IP, geographical location, and account information; the order information data includes the ticket purchase time, ticket type, ticket price, purchase quantity, and performance session; the payment information data includes the payment method, payment amount, and payment time; the user behavior data includes the ticket purchase frequency, refund times, and operation duration.
[0068] The method for performing preliminary processing on the acquired ticket-related data to obtain a comprehensive ticket feature data set includes:
[0069] Identifying and removing outliers in the user information data, order information data, payment information data, and user behavior data through a density clustering algorithm to obtain a user information feature data set, an order information feature data set, a payment information feature data set, and a user behavior feature data set;
[0070] Feature extraction is performed on the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset through TF-IDF and Laplace smoothing. Then, standard deviation normalization is carried out on the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset after feature extraction, converting them into a standard normal distribution with a mean of 0 and a standard deviation of 1, obtaining the normalized user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset;
[0071] The normalized user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset are fused through a weighted model to obtain a comprehensive ticket feature dataset.
[0072] The method of performing feature extraction on the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset through TF-IDF and Laplace smoothing includes:
[0073] For the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset, their term frequencies and inverse document frequencies are calculated respectively through the term frequency calculation formula and the inverse document frequency calculation formula; for example, for the user behavior dataset, the term frequency can be based on the user's behavior log, while the inverse document frequency reflects the scarcity of this behavior among all users.
[0074] The term frequency calculation formula is: where TF(ω j ,T i ) is the number of occurrences of the feature ω j in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset in the document T i ; ω j is a certain feature in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset; T i is a certain document in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset; in different datasets, T i represents a certain document or data record; for example, T 1 may be a user information entry, and T 2 is an order information entry, etc.; Co(ω j ,T i ) is the number of occurrences of the feature ω j in the document T iThe number of occurrences in; N is the total number of occurrences of all features in the document; j is the index of the feature; i is the index of the document;
[0075] The inverse document frequency calculation formula is: where D is the set of all documents; m is the total number of documents in the document set; n j is the number of documents containing the feature;
[0076] Calculate the TF-IDF value of each feature, generate a feature vector for each document, and the TF-IDF value calculation formula is: TF-IDF(ω j ,T i ,D) = TF(ω j ,T i ) · IDF(ω j ,D); where TF-IDF(ω j ,T i ,D) is the TF-IDF value of the feature ω j in the document T i ;
[0077] In the TF-IDF value, if a word feature ω j does not appear in a certain document T i (i.e., the word frequency is 0), then its word frequency is zero, resulting in its TF-IDF value also being zero. To avoid this situation, Laplace smoothing can be applied.
[0078] Smooth the word frequency of each feature through the Laplace smoothing formula, and the Laplace smoothing formula is: where TF sm (ω j ,T i ) is the word frequency value after Laplace smoothing; α d is the adaptive smoothing parameter, used to avoid the influence on the calculation when the number of occurrences of the feature ω j in the document T i is zero; V is the total number of all features in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset;
[0079] Dynamically design the adaptive smoothing parameter through the adaptive smoothing parameter formula, and the adaptive smoothing parameter formula is: where c 1 is the influence factor that controls the total number of occurrences N of all features in the document on the smoothing parameter; c 2 is the influence factor that controls the total number of all features V in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset on the smoothing parameter; c 3 is the constant factor that adjusts the overall smoothing effect;
[0080] For example, assume that the total number of occurrences N of all features in the control document is 50, and the total number V of all features in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset is 100. The influence factor c of the total number of occurrences N of all features in the control document on the smoothing parameter 1 is 0.5, and the influence factor c of the total number V of all features in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset on the smoothing parameter 2 is 0.3, and the constant factor c for adjusting the overall smoothing effect 3 is 0.2; then the adaptive smoothing parameter is
[0081]
[0082] By adjusting the word frequency of each feature through the above calculations and the Laplace smoothing formula, and finally combining the TF-IDF values of the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset, a multi-dimensional feature vector is generated.
[0083] The method for fusing the normalized user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset through a weighted model to obtain a comprehensive ticketing feature dataset includes:
[0084] Denote the user information feature dataset as F 1 , the order information feature dataset as F 2 , the payment information feature dataset as F 3 , and the user behavior feature dataset as F 4 ; the weighted model is: EPR = F 1 ·δ 1 + F 2 ·δ 2 + F 3 ·δ 3 + F 4 ·δ 4 ; where EPR is the comprehensive ticketing feature dataset; δ 1 is the weight coefficient of the user information feature dataset; δ 2 is the weight coefficient of the order information feature dataset; δ 3 is the weight coefficient of the payment information feature dataset; δ 4 is the weight coefficient of the user behavior feature dataset.
[0085] The preset rules include threshold rules, blacklist matching rules, and cross - domain rules; the threshold rule is based on preset numerical boundaries or ranges to check whether the feature values exceed these ranges. This rule identifies abnormal data that does not conform to the norm by restricting the absolute or relative values of the data.
[0086] For example, if the order amount exceeds a certain threshold (such as 1000 yuan), then the order is marked as abnormal; if a user places orders more than a preset number of times within a short period (such as buying more than 10 tickets in a day), then the user's behavior is considered abnormal; if the same user buys more than a certain number of tickets within a short period, then their behavior is considered abnormal.
[0087] The blacklist matching rule is used for anomaly detection based on a predefined blacklist. The blacklist may contain information such as disallowed users, IP addresses, devices, credit cards, etc. If the features in the order match the information in the blacklist, then the order is considered abnormal;
[0088] For example, if the IP address of an order appears in a known blacklist, it may be an order sent by a malicious user or a known fraud source; if the device ID is in the blacklist, it may indicate that the device is being used for fraudulent activities; if the credit card number used for payment belongs to the blacklist, then the payment behavior will be marked as abnormal;
[0089] The cross - domain rule involves the association between different data domains or systems, detecting abnormal situations that occur in different fields or different dimensions. This rule identifies potential fraud or improper behavior by examining the correlation of data, cross - domain anomalies in user behavior, etc.;
[0090] For example, a user logs in from different devices or different locations, but pays for multiple orders using the same account at the same time, which may be an act of theft or malicious behavior; a certain user's account has frequent login and information modification behaviors, but the payment behavior is very abnormal (such as quickly paying for a large - amount order), which may indicate that the account has been stolen. A user switches between multiple devices for payment, and the geographical locations of these devices are inconsistent with the historical activities of the account, which may be a manifestation of account theft.
[0091] The method for performing anomaly detection on the comprehensive ticket feature dataset according to the preset rules to obtain a suspected abnormal order feature dataset includes:
[0092] Performing anomaly detection on the comprehensive ticket feature dataset according to the threshold rule through an adaptive threshold formula. The adaptive threshold formula is: T u = β 1 ·log(N u + 1)+ β 2 ·avg(O u ) ; where T u is the adaptive threshold for user u; Nu The number of orders of user u within the past time window T; win O; u The order amount of user u within the past time window T; win avg(O u ) is the average value of the order amounts of user u within the past time window T; win β; 1 is the purchase behavior adjustment coefficient, which is used to reflect the influence of the user's purchase behavior within the time window on the adaptive threshold; 2 β is the average order amount adjustment coefficient, which is used to reflect the influence of the average order amount on the adaptive threshold; u is the index of the user;
[0093] According to the threshold rule, if the order amount or the number of orders of a certain user exceeds the adaptive threshold T of this user, u then this order is determined to be a suspected abnormal order;
[0094] Calculate the behavior score of the blacklist user through the blacklist score calculation formula. The blacklist score calculation formula is: where S b is the blacklist score of user u; is the purchase frequency of user u; is the volatility of the order amount of user u; σ(O u ) is the standard deviation of the order amount of user u; is the order speed; T ta is the average order interval time of user u; γ 1 is the purchase frequency adjustment coefficient, which is used to control the influence of the user's purchase frequency on the blacklist score; 2 γ is the order amount volatility adjustment coefficient, which is used to control the influence of the order amount volatility on the blacklist score; 3 γ is the order speed adjustment coefficient, which is used to control the influence of the order speed on the blacklist score;
[0095] According to the blacklist matching rule, preset the blacklist score threshold. When the blacklist score of user u exceeds the blacklist score threshold, then this user is determined to be a blacklist user;
[0096] According to the cross - domain rule, if user u switches different IP addresses or different devices within the past time window T win and pays for N orders with the same account, then it is determined that the user has abnormal payment behavior. u
[0097] The training method of the abnormal order prediction model includes:
[0098] Divide the dataset into a training set, a validation set, and a test set, and construct an abnormal order prediction model; the sample set is a subset of the dataset, and each sample set includes a historical suspected abnormal order feature dataset and the corresponding abnormal order evaluation label;
[0099] The input data of the model is the historical suspected abnormal order feature dataset, the output label of the model is the abnormal order evaluation label, and the abnormal order evaluation label is divided into two categories: abnormal orders and normal orders; use the sigmoid function as the activation function; the abnormal order prediction model is a gradient boosting decision tree model;
[0100] Initialize the abnormal order prediction model, and set the hyperparameters of the number of trees, the learning rate, and the depth of the tree; through k-fold cross-validation, use different combinations of hyperparameters to train the model, and use the results of cross-validation to tune the hyperparameters set initially, and select the best-performing combination of hyperparameters; use binary cross-entropy as the loss function to measure the error of the model prediction;
[0101] Train the abnormal order prediction model on the training set, and train a new decision tree based on the residuals of the current model in each iteration; according to the model performance feedback, adjust the hyperparameters of the model to optimize the model, retrain the model with the adjusted hyperparameters, and after the training process ends, select the optimal hyperparameters through cross-validation;
[0102] Stop when the training reaches the preset model complexity or convergence condition, and obtain the finally trained abnormal order prediction model; use the trained abnormal order prediction model to predict the current suspected abnormal order feature dataset to obtain the abnormal order evaluation label.
[0103] The method for adjusting the hyperparameters of the model to optimize the model includes:
[0104] S91. Randomly initialize the population, generate E individuals, and form the initial population G = {x 1 , x 2 ,..., x p ,..., x E}; where x p is the p-th individual in the population; E is the total number of individuals; p is the index of the individual, and p ∈ {1,..., E}; S92. Presuppose that the model has d hyperparameters, and each individual in the population represents a set of hyperparameter configurations; represent each individual as a d-dimensional vector: x p = (x p1 , x p2 ,..., x pd ); where x pd is the value of the d-th hyperparameter of the p-th individual;
[0105] S93. Measure the fitness of an individual through the binary cross - entropy loss function, and evaluate the current fitness f(x p ); Calculate the historical trend factor of each individual: Δf(x p ) = f(x p ) - f -1 (x p ); where, Δf(x p ) is the historical trend factor of each individual x p , representing the fitness improvement value of individual x p in this round of iteration; f -1 (x p ) is the fitness of individual x p in the previous round of iteration;
[0106] S94. Based on the fitness and the distance between individuals, calculate the love force through the love force formula to measure the intimacy between two individuals; The love force formula is: where, L(x p , x q ) is the love force between individual x p and individual x q ; ds(x p , x q ) is the distance between individual x p and individual x q , calculated by the Euclidean distance; f(x q ) is the fitness of individual x p ; Δf(x p ) is the historical trend factor of individual x q ; λ is the trend factor weight; Dynamically adjust the trend factor weight λ through the trend factor weight adjustment formula;
[0107] S95. Each individual updates its intimacy through the intimacy update formula according to the love force with other individuals; Intimacy reflects the relationship strength between individuals and affects their ability to exchange information and generate new solutions. The key to intimacy update lies in the interaction between individuals. The intimacy update formula is: where, Q(x p , t + 1) is the intimacy of individual x p at the (t + 1) - th iteration; Q(x p , t) is the intimacy of individual x p at the current t - th iteration;
[0108] S96. Preset an intimacy threshold. When two individuals x p and x qWhen the intimacy between them reaches the preset intimacy threshold, mating is performed through the mating formula to generate new individuals; the mating formula is: x new = a·x p +(1 - a)·x q ; where, x new is the new individual; a is a random factor that controls the mating degree of the two individuals and ranges between 0 and 1;
[0109] S97. Evaluate the fitness of the newly generated individual. If the fitness of the new individual is better, replace the original individual; otherwise, keep the original individual unchanged; preset the fitness threshold, and repeat S94 - S97 until the maximum number of iterations is reached or the preset fitness threshold is reached, and then stop. Finally, output the optimal hyperparameter configuration of the model, that is, the individual with the optimal fitness.
[0110] The trend factor weight adjustment formula is: where, λ′ is the adjusted trend factor weight; t is the current number of iterations; t max is the maximum number of iterations; b is an exponential factor that controls the adjustment rate of the trend factor weight.
[0111] The preset blacklist score threshold is set by the staff. Different blacklist scores are collected through the order anomaly detection terminal, and the average value of multiple blacklist scores is taken as the preset blacklist score threshold; similarly, the intimacy threshold and the fitness threshold are set.
[0112] In this embodiment, the importance of each feature can be quantified through the combination of TF - IDF; TF (term frequency) reflects the frequency of the feature in a single document, and IDF (inverse document frequency) suppresses the influence of common features by reducing the weights of features that appear frequently in multiple documents, thereby enhancing the attention to rare but important features; Laplace smoothing solves the situation where the term frequency is zero by adjusting the frequency of features, avoids calculation problems in the model due to some features not appearing in the document, can smooth the feature values, and enables the model to still maintain effective learning when facing rare features; by dynamically adjusting the adaptive smoothing parameter, the feature smoothing effect can be dynamically optimized according to the density of the document content or the distribution of features, improving the flexibility of the model;
[0113] The adaptive threshold formula can dynamically calculate the threshold for each user based on the number and amount of orders within a past time window, so as to adjust the detection sensitivity according to the historical behavior habits of different users. By setting the threshold in this dynamic, user behavior-driven manner, abnormal orders can be detected more accurately, avoiding misjudgments caused by fixed threshold settings; by dynamically adjusting the threshold and multi-dimensional abnormal detection rules, the false positive rate and false negative rate can be effectively reduced; the blacklist score can comprehensively evaluate the behavior of users. Combining various factors such as the purchase frequency, order amount volatility, and order speed of users, those users who frequently conduct abnormal transactions can be accurately identified. By calculating the blacklist score, the system can comprehensively analyze the behavior of users according to the score level and determine whether there is potential fraud behavior; the cross-domain rule can effectively detect the situation where a user uses the same account to make payments on different IP addresses or devices, which is a typical abnormal payment behavior. Through this cross-domain rule, risks such as identity theft and payment fraud by users can be discovered in a timely manner;
[0114] By randomly initializing the population and using the cross-entropy loss function to measure the fitness of individuals, different hyperparameter configurations can be efficiently explored to find the optimal hyperparameters suitable for the current task; this search method, by simulating the process of natural selection, can not only optimize a single hyperparameter but also adjust multiple hyperparameters simultaneously, thus greatly improving the overall performance of the model; through the historical trend factor and the love force formula, the optimization process of individuals and the intimacy between individuals can be dynamically evaluated, ensuring that during the optimization process, the intimacy between individuals is reasonably measured and adjusted, enabling the model to flexibly adapt to data changes during the iteration process; the dynamic adjustment of the trend factor weight makes the optimization process more meticulous. As the optimization process progresses, the weight factor can more reasonably affect the fitness evaluation of individuals, enhancing the accuracy of the search; by introducing an exponential factor to control the adjustment rate of the trend factor, the optimization process of the algorithm can dynamically adapt according to different task and data characteristics.
[0115] Embodiment 2
[0116] Please refer to Figure 2 As shown in the figure, for the parts not described in detail in this embodiment, refer to the description content of Embodiment 1. Provide a machine learning-based ticketing order anomaly detection system, including:
[0117] S1. Collect ticketing-related data;
[0118] S2. Perform preliminary processing on the collected ticketing-related data to obtain a comprehensive ticketing feature dataset;
[0119] S3. Perform anomaly detection on the comprehensive ticketing feature dataset according to preset rules to obtain a suspected abnormal order feature dataset;
[0120] S4. Train and obtain an abnormal order prediction model based on the suspected abnormal order feature dataset, predict the abnormal order evaluation label based on the abnormal order prediction model, and determine whether the ticket order is abnormal;
[0121] S5. If the ticket order is abnormal, the order abnormality detection terminal generates a warning message and sends a warning notice to relevant personnel.
[0122] Since the electronic device described in this embodiment is the electronic device adopted by the ticket order abnormality detection system based on machine learning in the embodiment of the present application, based on the ticket order abnormality detection system based on machine learning described in the embodiment of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiment of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device adopted by the ticket order abnormality detection system based on machine learning in the embodiment of the present application, it falls within the protection scope of the present application.
[0123] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formula are set by those skilled in the art according to the actual situation.
[0124] The above description is only a preferred embodiment of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for ordinary technical users in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. The machine learning-based ticket order anomaly detection system is characterized by: include: Data collection module, used to collect ticketing related data; The data preprocessing module is used to preliminarily process the collected ticket-related data to obtain a comprehensive ticket feature data set; The abnormal rule engine module is used to perform abnormal detection on the comprehensive ticket feature data set according to preset rules to obtain a suspected abnormal order feature data set; The abnormal order detection module is used to train an abnormal order prediction model based on the suspected abnormal order feature data set, obtain abnormal order evaluation labels based on the abnormal order prediction model, and determine whether the ticket order is abnormal; Abnormal warning module: if the ticket order is abnormal, the order abnormality detection terminal generates warning information and sends a warning notification to relevant personnel; each module is connected by wired and / or wireless means.
2. The machine learning-based ticket order anomaly detection system according to claim 1, characterized in that: The ticketing-related data includes user information data, order information data, payment information data and user behavior data; User information data includes user login IP, geographic location and account information; order information data includes purchase time, ticket type, ticket price, number of tickets purchased and number of sessions purchased; payment information data includes payment method, payment amount and payment time; user behavior data includes ticket purchase frequency, number of refunds and operation time.
3. The machine learning-based ticket order anomaly detection system according to claim 2, characterized in that: The method of performing preliminary processing on the collected ticket-related data to obtain a comprehensive ticket feature data set includes: Identify and remove outliers in user information data, order information data, payment information data, and user behavior data through a density clustering algorithm to obtain a user information feature data set, an order information feature data set, a payment information feature data set, and a user behavior feature data set; Perform feature extraction on the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset through TF-IDF and Laplace smoothing, and perform standard deviation normalization on the feature-extracted user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset, and convert them into a standard normal distribution with a mean of 0 and a standard deviation of 1, thereby obtaining the normalized user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset; The normalized user information feature dataset, order information feature dataset, payment information feature dataset and user behavior feature dataset are fused through a weighted model to obtain a comprehensive ticketing feature dataset.
4. The machine learning-based ticket order anomaly detection system according to claim 3, characterized in that: The method for extracting features from a user information feature data set, an order information feature data set, a payment information feature data set, and a user behavior feature data set by using TF-IDF and Laplace smoothing includes: For the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset, the word frequency and inverse document frequency are calculated using the word frequency calculation formula and the inverse document frequency calculation formula respectively; The formula for calculating word frequency is: Among them, TF(ω j ,T i ) is the feature ω in the user information feature dataset, order information feature dataset, payment information feature dataset and user behavior feature dataset j In the document T i The number of occurrences in ; ω j is a feature in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset; T i is a document in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset; Co(ω j ,T i ) is the characteristic ω j In the document T i N is the total number of occurrences of all features in the document; j is the index of the feature; i is the index of the document; The inverse document frequency calculation formula is: Where D is the set of all documents; m is the total number of documents in the document set; n j is the number of documents containing the feature; Calculate the TF-IDF value of each feature and generate a feature vector for each document. The TF-IDF value calculation formula is: TF-IDF(ω j ,T i ,D)=TF(ω j ,T i )·IDF(ω j ,D); Among them, TF-IDF(ω j ,T i ,D) is the characteristic ω j In the document T i TF-IDF value in; The word frequency of each feature is smoothed by the Laplace smoothing formula. The Laplace smoothing formula is: Among them, TF sm (ω j ,T i ) is the word frequency value after Laplace smoothing; α d is the adaptive smoothing parameter; V is the total number of all features in the user information feature dataset, order information feature dataset, payment information feature dataset, and user behavior feature dataset; The adaptive smoothing parameter is dynamically designed through the adaptive smoothing parameter formula, and the adaptive smoothing parameter formula is: Among them, c1 is the influence factor of the total number of occurrences N of all features in the control document on the smoothing parameter; c2 is the influence factor of the total number V of all features in the user information feature data set, order information feature data set, payment information feature data set and user behavior feature data set on the smoothing parameter; c3 is the constant factor that adjusts the overall smoothing effect; The word frequency of each feature is adjusted through the above calculation and Laplace smoothing formula, and finally a multidimensional feature vector is generated by combining the TF-IDF values of the user information feature dataset, order information feature dataset, payment information feature dataset and user behavior feature dataset.
5. The machine learning-based ticket order anomaly detection system according to claim 4, characterized in that: The method of fusing the normalized user information feature data set, order information feature data set, payment information feature data set and user behavior feature data set through a weighted model to obtain a comprehensive ticketing feature data set includes: The user information feature data set is denoted as F1, the order information feature data set is denoted as F2, the payment information feature data set is denoted as F3, and the user behavior feature data set is denoted as F4; the weighted model is: EPR=F1·δ1+F2·δ2+F3·δ3+F4·δ4; wherein EPR is the comprehensive ticketing feature data set; δ1 is the weight coefficient of the user information feature data set; δ2 is the weight coefficient of the order information feature data set; δ3 is the weight coefficient of the payment information feature data set; δ4 is the weight coefficient of the user behavior feature data set.
6. The machine learning-based ticket order anomaly detection system according to claim 5, characterized in that: The preset rules include threshold rules, blacklist matching rules and cross-domain rules.
7. The machine learning-based ticket order anomaly detection system according to claim 6, characterized in that: The method of performing anomaly detection on the comprehensive ticketing feature data set according to preset rules to obtain a suspected abnormal order feature data set includes: The adaptive threshold formula is used to detect anomalies in the comprehensive ticket feature data set according to the threshold rule. The adaptive threshold formula is: T u =β1·log(N u +1)+β2·avg(O u ), where T u is the adaptive threshold of user u; N u For users in the past T win The number of orders within the time window; u For user u past T win The order amount within the time window; avg(O u ) is the past T of user u win The average value of the order amount within the time window; β1 is the purchase behavior adjustment coefficient; β2 is the average order amount adjustment coefficient; u is the user's index; According to the threshold rule, if the order amount or number of orders of a user exceeds the user's adaptive threshold T u , then the order is determined to be a suspected abnormal order; The behavior score of blacklist users is calculated using the blacklist score calculation formula. The blacklist score calculation formula is: Among them, S b Score the blacklist of user u; is the purchase frequency of user u; is the volatility of the order amount of user u; σ(O u ) is the standard deviation of the order amount of user u; is the order speed; T ta is the average order interval time of user u; γ1 is the purchase frequency adjustment coefficient; γ2 is the order amount volatility adjustment coefficient; γ3 is the order speed adjustment coefficient; According to the blacklist matching rule, a blacklist score threshold is preset. When the blacklist score of user u exceeds the blacklist score threshold, the user is determined to be a blacklist user; According to the cross-domain rule, if user u has been in the past T win Switch to different IP addresses or different devices within the time window and pay N with the same account u If an order is placed, it is determined that the user has abnormal payment behavior.
8. The machine learning-based ticket order anomaly detection system according to claim 7, characterized in that: The training method of the abnormal order prediction model includes: The dataset is divided into training set, validation set and test set to build an abnormal order prediction model; the sample set is a subset of the dataset, and each sample set includes a historical suspected abnormal order feature dataset and the corresponding abnormal order evaluation label; The input data of the model is a historical suspected abnormal order feature data set, and the output label of the model is an abnormal order evaluation label, which is divided into two categories: abnormal orders and normal orders; the sigmoid function is used as the activation function; the abnormal order prediction model is a gradient boosting decision tree model; Initialize the abnormal order prediction model and set the number of trees, learning rate, and tree depth hyperparameters. Use k-fold cross validation to train the model using different hyperparameter combinations, and use the cross validation results to tune the initially set hyperparameters and select the best performing hyperparameter combination. Use binary cross entropy as the loss function to measure the error of the model prediction. The abnormal order prediction model is trained in the training set by training a new decision tree based on the residual of the current model in each iteration; the model hyperparameters are adjusted to optimize the model according to the model performance feedback, and the model is retrained with the adjusted hyperparameters. After the training process is completed, the optimal hyperparameters are selected through cross-validation; When the training reaches the preset model complexity or convergence condition, it stops and obtains the final trained abnormal order prediction model; the trained abnormal order prediction model is used to predict the current suspected abnormal order feature data set to obtain the abnormal order evaluation label.
9. The machine learning-based ticket order anomaly detection system according to claim 8, characterized in that: The method for adjusting the hyperparameters of the model to optimize the model includes: S91, randomly initialize the population, generate E individuals, and form the initial population G = {x1, x2, ..., x p ,...,x E }; where x p is the pth individual in the population; E is the total number of individuals; p is the index of the individual, p∈{1,...,E}; S92. The default model has d hyperparameters. Each individual in the population represents a set of hyperparameter configurations. Each individual is represented as a d-dimensional vector: x p =(x p1 ,x p2 ,...,x pd ), where x pd is the dth hyperparameter value of the pth individual; S93, measure the fitness of individuals through the binary cross entropy loss function, and evaluate the current fitness of each individual f(x p ); Calculate the historical trend factor for each individual: Δf(x p )=f(x p )-f -1 (x p ), where Δf(x p ) for each individual x p The historical trend factor represents the individual x p The fitness improvement value in this round of iteration; f -1 (x p ) is individual x p Fitness in the previous iteration; S94. Based on fitness and the distance between individuals, the love power is calculated by the love power formula to measure the intimacy between two individuals; the love power formula is: Among them, L(x p ,x q ) is individual x p and individual x q The power of love between; ds(x p ,x q ) is individual x p and individual x q The distance between them; f(x q ) is individual x q The fitness of Δf(x q ) is individual x q The historical trend factor; λ is the trend factor weight; the trend factor weight λ is dynamically adjusted through the trend factor weight adjustment formula; S95. Each individual updates its intimacy according to the love power with other individuals through the intimacy update formula; the intimacy update formula is: Among them, Q(x p ,t+1) is individual x p The intimacy at the t+1th iteration; Q(x p ,t) is the individual x p Intimacy at the current t-th iteration; S96, preset intimacy threshold, when two individuals x p and x q When the intimacy between them reaches the preset intimacy threshold, they mate through the mating formula to generate new individuals; the mating formula is: new =a·x p +(1-a)·x q ; where x new is a new individual; a is a random factor that controls the degree of mating between two individuals, between 0 and 1; S97, evaluate the fitness of the newly generated individual. If the new individual has better fitness, replace the original individual; otherwise, keep the original individual unchanged; preset the fitness threshold, repeat S94-S97 until the maximum number of iterations is reached or the preset fitness threshold is reached, and finally output the optimal hyperparameter configuration of the model, that is, the individual with the optimal fitness.
10. The machine learning-based ticket order anomaly detection system according to claim 9, characterized in that: The trend factor weight adjustment formula is: Among them, λ′ is the adjusted trend factor weight; t is the current iteration number; t max is the maximum number of iterations; b is the exponential factor that controls the rate of adjustment of the trend factor weight.
Citation Information
Patent Citations
Abnormal order detection and early warning method, system and device and storage medium
CN116934418A
Cited By
Distributed system and device for payment state monitoring and exception handling, and storage medium
CN120996796A
Big data-based movie ticket selling intelligent operation management system
CN122198583A