Automatic commission tracing checking method and device and computer readable storage medium
Through automated commission traceability inspection methods, including data preprocessing, feature analysis and machine learning model training based on gradient-enhanced decision tree, the problem of inefficiency of manual commission traceability on large-scale data sets is solved, and more efficient and accurate commission traceability detection is achieved.
Patent Information
- Application Number
- CN202510138860.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, artificial commission traceability is relatively inefficient on large-scale data sets, and has problems such as low efficiency and error proneness, which is difficult to meet current market demand, especially when dealing with large-scale data and high-frequency transactions.
Automatic commission traceability inspection methods are adopted, including obtaining historical commission data from data sources and preprocessing, using multiple statistical analysis methods for feature analysis, determining target feature information, using machine learning algorithms based on gradient enhancement decision tree to build prediction models, training and optimization, and finally predicting and anomaly detection of the current commission data.
It improves the efficiency and accuracy of commission traceability, reduces the burden of manual auditing, enhances the coverage and efficiency of detection, and can quickly locate potential commission traceability problems on large-scale data sets.
Smart Images

Figure CN120069951A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent data processing and analysis. Specifically, it relates to an automated commission traceability inspection method, device, computer-readable storage medium, and electronic device. Background Art
[0002] In the field of the operator industry, commission traceability involves re-calculating the settled billing periods to accurately trace the commissions generated during historical billing periods when adjusting expenses. This process is crucial for the rationality and accuracy of settlements. Under the traditional management mode, the detection process of commission traceability mainly relies on manual operations, comparing data between systems, and auditing the traced data, which has problems such as low efficiency and being prone to errors. Moreover, with the expansion of the market scale and the increase in business volume, the manual processing method has become difficult to meet the current needs. Especially when dealing with large-scale data and high-frequency transactions, the issues of efficiency and accuracy have become increasingly prominent.
[0003] Traditional manual commission traceability faces many challenges. Processing massive amounts of data not only requires a large amount of manpower and time but also easily introduces human errors, affecting the accuracy of comparison results. Moreover, due to the limited speed of manual data processing, it is difficult to compare and check data in real-time or quickly. In addition, frequent commission traceability inspections consume a large amount of manpower and computing resources, increasing costs. All in all, manual commission traceability is difficult to quickly adapt and adjust, reducing the flexibility and efficiency of traceability. This results in low efficiency in commission traceability on large-scale data sets. Summary of the Invention
[0004] The main purpose of this application is to provide an automated commission traceability inspection method, device, computer-readable storage medium, and electronic device to at least solve the problem of low efficiency in commission traceability of manual commission traceability in large-scale data sets in the prior art.
[0005] To achieve the above object, according to one aspect of the present application, an automated commission traceability inspection method is provided, including: obtaining historical commission data from a data source and preprocessing the historical commission data to obtain preprocessed historical commission data, wherein the historical commission data includes various feature information related to commission traceability; performing feature analysis on the preprocessed historical commission data by using various statistical analysis methods to obtain a comprehensive relevance ranking of the various feature information, and determining the feature information within a first preset range of the comprehensive relevance ranking as target feature information; constructing an initial commission traceability prediction model by using a machine learning algorithm based on a gradient boosting decision tree, inputting the target feature information into the initial commission traceability prediction model for training and optimization to obtain a target commission traceability prediction model; inputting current commission data into the target commission traceability prediction model for prediction to obtain a prediction result, marking the data records whose prediction results are not within a second preset range as abnormal data records, and outputting the abnormal data records.
[0006] Optionally, performing feature analysis on the preprocessed historical commission data by using various statistical analysis methods to obtain a comprehensive relevance ranking of the various feature information, and determining the feature information within a first preset range of the comprehensive relevance ranking as target feature information includes: performing feature analysis on the preprocessed historical commission data by using various statistical analysis methods to obtain multiple relevance degrees of each feature information with commission traceability, arranging the values of the multiple relevance degrees obtained by each statistical analysis method from large to small to obtain a relevance ranking column vector corresponding to each statistical analysis method; performing weighted construction on the relevance ranking column vector corresponding to each statistical analysis method to construct a feature comprehensive screening model, solving the feature comprehensive screening model to obtain a comprehensive relevance column vector of the feature information; and determining the feature information within the first preset range of the comprehensive relevance column vector as target feature information.
[0007] Optionally, various statistical analysis methods are used to perform feature analysis on the preprocessed historical commission data to obtain multiple degrees of correlation between each of the feature information and commission tracing, and the values of the multiple degrees of correlation obtained by each of the statistical analysis methods are arranged from largest to smallest to obtain a correlation ranking column vector corresponding to each of the statistical analysis methods, including: calculating multiple first degrees of correlation between each of the feature information and the commission tracing variable by using the correlation coefficient matrix method, and arranging the values of the multiple first degrees of correlation from largest to smallest to obtain a first correlation ranking column vector; calculating a second degree of correlation between each of the feature information and the commission tracing variable by using the distance correlation coefficient, and arranging the values of the multiple second degrees of correlation from largest to smallest to obtain a second correlation ranking column vector; calculating a third degree of correlation between each of the feature information and the commission tracing variable by using the recursive feature elimination method, and arranging the values of the multiple third degrees of correlation from largest to smallest to obtain a third correlation ranking column vector; calculating a fourth degree of correlation between each of the feature information and the commission tracing variable by using the grey relational analysis method, and arranging the values of the multiple fourth degrees of correlation from largest to smallest to obtain a fourth correlation ranking column vector.
[0008] Optionally, a feature comprehensive screening model is constructed by weighting the correlation ranking column vector corresponding to each of the statistical analysis methods, and the feature comprehensive screening model is solved to obtain a comprehensive correlation column vector of the feature information, including: constructing a feature comprehensive screening model by weighting the correlation ranking column vector where Y is the comprehensive correlation column vector, and β i represents the weight assigned to the i-th statistical analysis method; y i represents the correlation ranking column vector obtained by the i-th statistical analysis method; by solving the feature comprehensive screening model, the comprehensive correlation column vector arranged from highest to lowest in terms of correlation is obtained.
[0009] Optionally, use a machine learning algorithm based on gradient boosting decision trees to construct an initial commission retrospective prediction model, input the target feature information into the initial commission retrospective prediction model for training and optimization to obtain a target commission retrospective prediction model, including: a first construction step of using the LightGBM algorithm to construct the initial commission retrospective prediction model; a first calculation step of inputting the target feature information into the initial commission retrospective prediction model for training and calculating the loss function of the initial commission retrospective prediction model, where the loss function is used to quantify the difference between the predicted value and the true value of the initial commission retrospective prediction model; a second calculation step of calculating the negative gradient of the loss function with respect to the predicted value for each sample data point, where the negative gradient represents the optimization direction of the initial commission retrospective prediction model at the sample data point; a second construction step of constructing a new decision tree according to the negative gradient, where the goal of the new decision tree is to fit the negative gradient to reduce the loss function; an update step of adding the new decision tree to the initial commission retrospective prediction model and updating the predicted value of the initial commission retrospective prediction model; repeating the first calculation step, the second calculation step, the second construction step and the update step until a preset stop condition is reached to obtain the target commission retrospective prediction model.
[0010] Optionally, use a machine learning algorithm based on gradient boosting decision trees to construct an initial commission retrospective prediction model, input the target feature information into the initial commission retrospective prediction model for training and optimization to obtain a target commission retrospective prediction model, and further include: using the Leaf-wise decision tree growth strategy and the histogram-based gradient boosting method to optimize the target commission retrospective prediction model.
[0011] Optionally, after obtaining the target commission retrospective prediction model, the method further includes: adopting the coefficient of determination root mean square error mean absolute error and mean absolute percentage error to evaluate the accuracy of the target commission retrospective prediction model, where n represents the number of samples, y i represents the i-th measured value, represents the i-th predicted value, represents the mean of the measured values.
[0012] According to another aspect of the present application, an automated commission traceability inspection device is provided, including: an acquisition unit configured to acquire historical commission data from a data source and preprocess the historical commission data to obtain preprocessed historical commission data, wherein the historical commission data includes various feature information related to commission traceability; a determination unit configured to perform feature analysis on the preprocessed historical commission data by using various statistical analysis methods to obtain a comprehensive relevance ranking of the various feature information, and determine the feature information within a first preset range of the comprehensive relevance ranking as target feature information; a training and optimization unit configured to construct an initial commission traceability prediction model by using a machine learning algorithm based on gradient boosting decision trees, input the target feature information into the initial commission traceability prediction model for training and optimization to obtain a target commission traceability prediction model; and an output unit configured to input current commission data into the target commission traceability prediction model for prediction to obtain a prediction result, mark the data records whose prediction results are not within a second preset range as abnormal data records, and output the abnormal data records.
[0013] According to still another aspect of the present application, a computer-readable storage medium is provided, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the automated commission traceability inspection methods.
[0014] According to yet another aspect of the present application, an electronic device is provided, including: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the automated commission traceability inspection methods.
[0015] Applying the technical solution of the present application, first, historical commission data is obtained from a data source, and the historical commission data is preprocessed to obtain preprocessed historical commission data, where the historical commission data includes various feature information related to commission tracing; then, various statistical analysis methods are used to perform feature analysis on the preprocessed historical commission data to obtain a comprehensive relevance ranking of various feature information, and the feature information within the first preset range of the comprehensive relevance ranking is determined as target feature information; a machine learning algorithm based on gradient boosting decision tree is used to construct an initial commission tracing prediction model, and the target feature information is input into the initial commission tracing prediction model for training and optimization to obtain a target commission tracing prediction model; finally, the current commission data is input into the target commission tracing prediction model for prediction to obtain a prediction result, and the data records whose prediction results are not within the second preset range are marked as abnormal data records, and the abnormal data records are output. This solution solves the problem of low efficiency of manual commission tracing in large-scale data sets in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0017] Figure 1 The hardware structure block diagram of a mobile terminal showing an automated commission tracing inspection method provided in an embodiment of this application is shown;
[0018] Figure 2 The flowchart of an automated commission tracing inspection method provided in an embodiment of this application is shown;
[0019] Figure 3 The schematic diagram of the Level-wise tree growth strategy of an automated commission tracing inspection method provided in an embodiment of this application is shown;
[0020] Figure 4 The schematic diagram of the Leaf-wise tree growth strategy of an automated commission tracing inspection method provided in an embodiment of this application is shown;
[0021] Figure 5 The schematic diagram of finding the optimal splitting point of an automated commission tracing inspection method provided in an embodiment of this application is shown;
[0022] Figure 6 The structure block diagram of an automated commission tracing inspection device provided in an embodiment of this application is shown.
[0023] Wherein, the above-mentioned accompanying drawings include the following reference numerals:
[0024] 102, Processor; 104, Memory; 106, Transmission Device; 108, Input / Output Device. Detailed Implementation Manner
[0025] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0026] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0028] As introduced in the background art, manual commission tracing in the prior art requires a large amount of manpower and time, and the accuracy is relatively low. To solve the problem of low efficiency of commission tracing for manual commission tracing on a large-scale data set, the embodiments of the present application provide an automated commission tracing inspection method, device, computer-readable storage medium, and electronic device.
[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention.
[0030] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is the hardware structure block diagram of a mobile terminal of an automated commission tracing inspection method according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1Only one processor 102 is shown (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only illustrative and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in, or have a different configuration from Figure 1 shown.
[0031] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the display method of device information in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0032] In this embodiment, an automated commission traceability inspection method running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0033] Figure 2 is a flowchart diagram of an automated commission traceability inspection method according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0034] Step S201: Obtain historical commission data from a data source and preprocess the historical commission data to obtain preprocessed historical commission data. Among them, the historical commission data includes various feature information related to commission traceability.
[0035] Specifically, the data source refers to a system or database that stores or generates historical commission data. In the context of an operator, it usually includes a partner settlement subsystem, a Customer Relationship Management (CRM) system, a billing system, etc. These systems record various transaction and settlement information related to commission traceability. For example, the settlement record table will contain key field information such as settlement amount, settlement time, order ID, sales product ID, expense item ID, settlement rule, settlement ratio, etc. These are all components of the historical commission data.
[0036] Historical commission data refers to commission settlement records generated over a past period (such as the 6 months or 1 year before the traceability period), including settlement amount, settlement time, sales product type, expense item type, business rules, settlement ratio, etc. These data are crucial for model training because they contain various situations and patterns that may be encountered in commission traceability.
[0037] Data preprocessing is a crucial step in a machine learning project. Its goal is to clean and organize the original commission traceability data to ensure data quality. Data preprocessing mainly includes data transformation and data deduplication. Among them, data transformation is to standardize fields such as dates in the original data to ensure that all formats are unified into a standard format for subsequent data analysis and processing. Data deduplication is that duplicate records will cause bias in model training and affect the prediction ability of the model. By performing data deduplication operations, duplicate records are deleted to ensure the uniqueness and accuracy of the data.
[0038] Feature information refers to various fields or variables related to commission traceability that are retained during the data preprocessing process. These features can be numerical, categorical, or time series, and are the basis for subsequent feature analysis and model training. Feature information related to commission traceability is a feature that affects commission traceability, including settlement amount, settlement time, sales product id, expense item id, settlement rule, settlement ratio.
[0039] The preprocessed historical commission data not only reduces data noise, improves data quality, but also ensures data consistency and availability, providing a reliable input for subsequent feature analysis and model training. The purpose of this stage is to prepare a clean and uniformly formatted dataset for the subsequent machine learning process, enabling the model to effectively learn and predict based on high-quality data.
[0040] In step S202, various statistical analysis methods are used to perform feature analysis on the preprocessed historical commission data, obtaining a comprehensive relevance ranking of various pieces of the above-mentioned feature information, and determining the feature information with the comprehensive relevance ranking within the first preset range as the target feature information;
[0041] Specifically, on the preprocessed historical commission data, a series of statistical analysis methods are used to identify and analyze the relationship strength between each piece of feature information and the commission traceability target. The series of statistical analysis methods includes the correlation coefficient matrix method, distance correlation coefficient, recursive feature elimination method, and grey relational analysis method. Each method has its specific statistic and advantages and disadvantages, and can be used alone or in combination to evaluate the relevance of features.
[0042] Each statistical analysis method assigns a relevance ranking to the feature information, and the relevance ranking reflects the importance of the feature information for the commission traceability prediction task. The feature relevances obtained through different statistical analysis methods are weighted and synthesized to obtain a comprehensive relevance ranking. Each method may be assigned different weights according to its applicability and reliability in a specific scenario, so as to ensure that the final ranking can comprehensively reflect the importance of features.
[0043] From the comprehensive relevance ranking, the feature information ranked within the first preset range is selected as the target feature information. In the embodiments of this application, the first preset range is the top 20 ranked feature information. These feature information will be used in the model training stage to construct a model for predicting the accuracy of commission traceability. Through this optimization strategy, the complexity of the model can be reduced, the training efficiency and prediction performance of the model can be improved, and at the same time, the negative impact of irrelevant or redundant features on the model can be avoided.
[0044] Performing feature analysis using various statistical analysis methods, obtaining a comprehensive relevance ranking, and determining the target feature information are key steps to ensure that the machine learning model can effectively learn from historical commission data. This helps to construct a more powerful and accurate commission traceability prediction model, thereby improving the efficiency and reliability of automated commission traceability checks.
[0045] In step S203, a machine learning algorithm based on gradient boosting decision trees is used to construct an initial commission traceability prediction model, and the above-mentioned target feature information is input into the initial commission traceability prediction model for training and optimization to obtain a target commission traceability prediction model;
[0046] Specifically, the model construction phase involves machine learning algorithms. In the embodiments of the present application, the LightGBM algorithm is used to construct an initial commission traceability prediction model. LightGBM is an efficient distributed framework based on Gradient Boosting Decision Tree (GBDT). Its main features are efficient model training and prediction capabilities, suitable for processing large-scale data sets. To improve the training speed and efficiency of traditional GBDT, LightGBM introduces some key improvement methods during the decision tree splitting process. LightGBM is an efficient variant of GBDT, which realizes fast training of large-scale data sets through optimization techniques such as Leaf-wise tree growth strategy and histogram-based gradient boosting.
[0047] The obtained target feature information is input into the initial commission traceability prediction model for training. The purpose is to let the algorithm learn the patterns or rules of commission traceability accuracy through historical data. The training process involves inputting the target feature information together with the results of commission traceability in historical commission data (such as whether it is accurate, whether there are problems) into the model. The LightGBM model will, through an iterative process, learn how to convert the input features into prediction results of commission traceability. In each iteration, the model will adjust the structure and parameters of the decision tree according to the current prediction error to gradually reduce the prediction error.
[0048] Model optimization refers to the process of adjusting model parameters during training to improve model performance, including adjusting the learning rate, the maximum depth of the tree, the minimum number of samples in leaf nodes, etc. The goal of optimization is usually to minimize the loss function of the model (such as mean square error, logarithmic loss, etc.) through cross-validation (such as using a validation set), while avoiding model overfitting (i.e., performing well on training data but poorly on new data). In the embodiments of the present application, model optimization will continue until the model reaches a predetermined accuracy standard or stop condition.
[0049] By using a machine learning algorithm based on gradient boosting decision tree to construct and optimize the model, we can obtain a model that can efficiently and accurately predict the accuracy of commission traceability. This process not only reduces the need for manual intervention but also improves the efficiency and accuracy of commission traceability detection.
[0050] In step S204, the current commission data is input into the above-mentioned target commission traceability prediction model for prediction to obtain a prediction result. The data records whose prediction results are not within the second preset range are marked as abnormal data records, and the above-mentioned abnormal data records are output.
[0051] Specifically, after the model construction and optimization phase are completed, the model already has the ability to predict the accuracy of commission retrospective based on the selected features. At this time, the latest commission data, that is, the commission records for the current accounting period, need to be input into this already trained target commission retrospective prediction model, and these latest commission data contain the same feature information as the historical data.
[0052] After receiving the latest commission data, the target commission retrospective prediction model will make predictions according to the learned patterns and rules, and output a prediction result representing the accuracy of commission retrospective. The prediction result can be a continuous value (such as the expected amount or accuracy score of commission retrospective), or a classification label (such as whether the commission retrospective is accurate, whether there are anomalies, etc.).
[0053] To identify abnormal commission data, a "normal" range needs to be set as the second preset range, which is determined based on the prediction distribution of the model and business requirements. For example, if the model predicts the accuracy score of commission retrospective, a threshold can be set. For instance, a score between 90% - 100% is considered normal, and a score below 90% is regarded as abnormal.
[0054] The prediction result of the model is compared with the second preset range. If the prediction result of a certain data record is not within this "normal" range, then this data record will be marked as an abnormal data record. Abnormal markings are usually used to identify those commission records that the model determines have potential problems or anomalies. Finally, all data records marked as abnormal are output for manual review or further automated processing.
[0055] By inputting the latest commission data into the target commission retrospective prediction model for prediction, marking and outputting abnormal data records, the automated commission retrospective inspection method can quickly locate potential commission retrospective problems in large datasets, which is crucial for timely correcting errors, improving settlement accuracy and credibility. This method not only reduces the burden of manual review but also improves the detection coverage and efficiency.
[0056] In this embodiment, first, historical commission data is obtained from a data source, and the historical commission data is preprocessed to obtain the preprocessed historical commission data, where the historical commission data includes various feature information related to commission tracing; then, various statistical analysis methods are used to perform feature analysis on the preprocessed historical commission data to obtain a comprehensive relevance ranking of various feature information, and the feature information with the comprehensive relevance ranking within a first preset range is determined as the target feature information; a machine learning algorithm based on gradient boosting decision tree is used to construct an initial commission tracing prediction model, and the target feature information is input into the initial commission tracing prediction model for training and optimization to obtain a target commission tracing prediction model; finally, the current commission data is input into the target commission tracing prediction model for prediction to obtain a prediction result, and the data records whose prediction results are not within a second preset range are marked as abnormal data records, and the abnormal data records are output, thereby solving the problem of low efficiency of manual commission tracing in large-scale data sets in the prior art.
[0057] In one embodiment of the present application, various statistical analysis methods are used to perform feature analysis on the above-mentioned preprocessed historical commission data to obtain a comprehensive relevance ranking of various above-mentioned feature information, and the above-mentioned feature information with the comprehensive relevance ranking within a first preset range is determined as the target feature information, including: using various statistical analysis methods to perform feature analysis on the above-mentioned preprocessed historical commission data to obtain multiple relevance degrees of each above-mentioned feature information with commission tracing, and arranging the values of the multiple above-mentioned relevance degrees obtained by each above-mentioned statistical analysis method from large to small to obtain a relevance ranking column vector corresponding to each above-mentioned statistical analysis method; performing weighted construction on the above-mentioned relevance ranking column vector corresponding to each above-mentioned statistical analysis method to construct a feature comprehensive screening model, solving the above-mentioned feature comprehensive screening model to obtain a comprehensive relevance column vector of the above-mentioned feature information; determining the above-mentioned feature information within the above-mentioned first preset range in the above-mentioned comprehensive relevance column vector as the target feature information.
[0058] Specifically, first, feature analysis is performed on the preprocessed historical commission data, and this step aims to quantify the degree of association between each feature information and the accuracy of commission tracing. Multiple statistical analysis methods are simultaneously applied to the historical commission data, including the correlation coefficient matrix method, distance correlation coefficient, recursive feature elimination method, and grey relational analysis method. Each method will generate a relevance with commission tracing for each feature information, and the relevance reflects the importance of the feature information in predicting the accuracy of commission tracing.
[0059] The relevance values obtained by each statistical analysis method are sorted from large to small to form a relevance ranking column vector. This represents the ranking of the importance of different feature information to the commission tracing prediction task according to each method. Each relevance ranking column vector is an ordered list, and each feature information corresponds to a ranking. The higher the ranking, the more important the feature information is considered to be for the prediction task.
[0060] The feature comprehensive screening model is a model used to integrate the feature importance evaluations obtained by different statistical analysis methods. A weight is assigned to the relevance ranking column vector corresponding to each statistical analysis method, and the size of the weight reflects the relative importance of the method in the decision-making process. The weight setting can be based on the reliability of the method itself and its applicability in a specific scenario, or it can be automatically determined by methods such as cross-validation. Through the weighted summation method, the relevance ranking column vectors corresponding to each statistical analysis method are merged to obtain a comprehensive relevance column vector. The construction process of this column vector takes into account the evaluation results of all methods, and adjusts the degree of influence of different methods on the final result through weights.
[0061] The process of solving the feature comprehensive screening model is essentially to calculate the weighted importance of each feature under all method evaluations, so as to obtain a comprehensive relevance column vector of the feature information, that is, to obtain a list containing the comprehensive evaluation results of all features. According to the comprehensive relevance column vector, the feature information ranked in the first preset range is determined as the target feature information. The first preset range usually refers to the features ranked at the top of the comprehensive relevance. In the embodiment of the present application, the feature information ranked in the top 20 in the comprehensive relevance column vector is determined as the target feature information. These feature information are considered to have a significant impact on the prediction effect of commission tracing, and are therefore selected for subsequent model training and prediction.
[0062] The process of selecting target feature information is a strategy in feature selection, which aims to streamline model input while maintaining or improving prediction accuracy. By screening out the features most relevant to the target, the complexity of the model can be reduced, the training efficiency can be improved, and the risk of overfitting can be reduced.
[0063] In summary, this process uses a variety of statistical analysis methods to evaluate the characteristics of the data, and then combines these evaluation results to determine the most relevant and effective feature set. These target feature information will then be used to build and optimize the commission tracing prediction model, thereby improving the model's predictive performance and reliability, and providing a solid foundation for the accuracy of automated detection of commission tracing.
[0064] In a specific embodiment, multiple statistical analysis methods are used to perform feature analysis on the above-mentioned preprocessed historical commission data to obtain multiple correlations between each of the above-mentioned feature information and commission tracing, and the values of the multiple correlations obtained by each of the above-mentioned statistical analysis methods are arranged from largest to smallest to obtain a correlation ranking column vector corresponding to each of the above-mentioned statistical analysis methods, including: calculating multiple first correlations between each of the above-mentioned feature information and the commission tracing variable by using the correlation coefficient matrix method, and arranging the values of the multiple first correlations from largest to smallest to obtain a first correlation ranking column vector; calculating a second correlation between each of the above-mentioned feature information and the commission tracing variable by using the distance correlation coefficient, and arranging the values of the multiple second correlations from largest to smallest to obtain a second correlation ranking column vector; calculating a third correlation between each of the above-mentioned feature information and the commission tracing variable by using the recursive feature elimination method, and arranging the values of the multiple third correlations from largest to smallest to obtain a third correlation ranking column vector; calculating a fourth correlation between each of the above-mentioned feature information and the commission tracing variable by using the grey relational analysis method, and arranging the values of the multiple fourth correlations from largest to smallest to obtain a fourth correlation ranking column vector.
[0065] Specifically, the correlation coefficient matrix method is used to calculate the correlation between each feature information and the commission tracing variable, that is, the first correlation. The correlation coefficient matrix method includes multiple specific correlation metrics. The correlation coefficient matrix methods that can be used in the embodiments of the present application include three typical correlation coefficient algorithms: Pearson correlation coefficient, Spearman rank correlation coefficient, and Kendall rank correlation coefficient. They are all used to represent the trend direction and degree of change of two variables. Taking the mathematical model of the Pearson correlation coefficient algorithm as an example for introduction, its most application scenarios are to describe the measurement of linear correlation between two samples, and the value range is [-1, 1]. Let (X, Y) be a two-dimensional random variable vector, and the formula for the Pearson correlation coefficient is
[0066] where r is the Pearson correlation coefficient, representing the correlation between variable X and variable Y, which is obtained by dividing the covariance of two continuous variables (X, Y) by the product of their respective standard deviations. A negative number indicates negative correlation, and a positive number indicates positive correlation. n is the number of sample points, and X i 、Y i are the values of the i-th sample point on variables X and Y respectively, are the means of X and Y respectively. Under the premise of significance, the stronger the correlation, the larger the absolute value. An absolute value of 0 indicates no linear relationship; an absolute value of 1 indicates a perfect linear correlation.
[0067] After obtaining the first relevance of all feature information, sort the first relevance corresponding to all feature information from largest to smallest, so as to obtain the first relevance ranking column vector. In this column vector, features with high relevance are ranked in the front, indicating that these features have a more significant impact on the commission traceability variable.
[0068] The distance correlation coefficient is used to calculate the correlation between each feature information and the commission traceability variable, that is, the second relevance. The distance correlation coefficient method is mainly used to solve the problem of non-linear correlation. It is a new measure to describe the independence between any variables. Different from the classical correlation coefficient, especially it can overcome the deficiencies of the Pearson correlation coefficient. For example, in some cases, even if the Pearson correlation coefficient is 0, it cannot be concluded that two variables are independent (they may be non-linearly correlated), but when the distance correlation coefficient is 0, it can indicate that two variables are independent.
[0069] Judge the independence of random variables X and Y through the distance correlation coefficient. The formula for the distance correlation coefficient is Divide the distance covariance of X and Y by the product of the distance standard deviations to obtain the distance correlation coefficient dcor(X,Y). When dcor(X,Y) = 0, it means that the two variables are independent of each other; the larger dcor(X,Y) = 0 is, the stronger the correlation between the two variables is.
[0070] After obtaining the second relevance of all feature information, sort the second relevance corresponding to all feature information from largest to smallest, so as to obtain the second relevance ranking column vector. In this column vector, features with high relevance are ranked in the front, indicating that these features have a more significant impact on the commission traceability variable.
[0071] The main idea of Recursive Feature Elimination (RFE) is to repeatedly construct models (such as support vector machines or regression models), select the best (or worst) features based on the coefficients, and then repeat this process on the remaining features until all features are traversed. In this process, the order in which features are eliminated is the ranking of the features. Therefore, this is a greedy algorithm for finding the optimal feature subset.
[0072] The stability of recursive feature elimination depends to a large extent on the underlying model used during iteration. For example, if ordinary regression (unregularized regression, which is unstable) is used in RFE, then RFE is unstable; if Lasso (Lasso-regularized regression is stable) is used, then RFE is stable. Therefore, in the embodiments of the present application, stability selection is first performed on RFE algorithmically, and then feature regression with Lasso regularization is selected. Recursive feature elimination is a feature selection method that determines the third relevance by repeatedly training the model and removing the least influential features. In each round, the model is trained, and then the least important features are removed based on the evaluation of feature importance until the number of remaining features meets the preset requirements. Through multiple iterations, a feature ranking can be obtained, where the features that contribute more to the model's prediction ability are ranked at the forefront, forming the third relevance ranking column vector.
[0073] The grey relational analysis method is used to calculate the correlation between each feature information and the commission traceability variable, that is, the fourth relevance. This method is mainly used to evaluate the correlation between two sequences and is especially suitable for dealing with the dynamic correlation problem between features. By calculating the grey relational degree between the feature sequence and the commission traceability variable sequence, it can be determined which features have the most consistent change trend with the commission traceability variable over time. Sorting the fourth relevance from largest to smallest gives the fourth relevance ranking column vector.
[0074] Specifically, the measure of the correlation magnitude between factors of two systems changing over time or different objects is called the correlation degree. In the process of system development, if the change trends of two factors are consistent and the degree of synchronous change is high, the correlation degree between the two is high; otherwise, it is low. Therefore, the grey relational analysis method is a method for measuring the correlation degree of factors based on the similarity or dissimilarity of the development trends between factors (i.e., the "grey relational degree").
[0075] Due to the different attributes and measurement units of each index, in order to establish a unified evaluation standard, it is necessary to perform normalization processing on them. The formula is where, X i (k) represents the sample data value of the i-th variable at the k-th time point, represents the average value of this sequence. The normalized sequence value X i (k) will reflect the change trend of the variable relative to its average value, eliminating the influence of dimension and magnitude.
[0076] Let the feature information value be the reference sequence X 0 = x 0 (k), k = 1, 2,..., n, and the feature sequence X i = x i (k), k = 1, 2,..., n, i = 1, 2,..., m, then X 0 and Xi The grey relational degree r(X 0 , X i ) is defined as shown in the following mathematical formula:
[0077]
[0078] Among them, ρ is the resolution coefficient, and ρ ∈ [0, 1]. Its role is to improve the difference significance between the correlation coefficients. Usually, ρ = 0.5.
[0079] After obtaining the fourth correlation degree of all feature information, sort the fourth correlation degrees corresponding to all feature information from largest to smallest, so as to obtain the fourth correlation degree ranking column vector.
[0080] The correlation degree ranking column vectors obtained by each statistical analysis method reflect the feature importance evaluated from different perspectives. The construction process of these column vectors not only helps to identify the features highly correlated with commission tracing, but also provides a basis for the subsequent feature comprehensive screening model, that is, by weight allocation, the importance evaluations from different perspectives are fused to determine the final target feature information. This series of statistical analysis methods and evaluation processes ensure the comprehensiveness and accuracy of feature selection, thereby improving the performance of the subsequent commission tracing prediction model.
[0081] In a specific another embodiment, a feature comprehensive screening model is constructed by weighting the above-mentioned correlation degree ranking column vectors corresponding to each of the above statistical analysis methods, and the above-mentioned feature comprehensive screening model is solved to obtain the comprehensive correlation degree column vector of the above-mentioned feature information, including: constructing a feature comprehensive screening model by weighting the above-mentioned correlation degree ranking column vectors Among them, Y is the above-mentioned comprehensive correlation degree column vector, and β i represents the weight assigned to the i-th above-mentioned statistical analysis method; y i represents the correlation degree ranking column vector obtained by the i-th above-mentioned statistical analysis method; by solving the feature comprehensive screening model, the above-mentioned comprehensive correlation degree column vector arranged from high to low in correlation degree is obtained.
[0082] Specifically, the construction of the feature screening model is based on the correlation degree ranking column vectors obtained by the following several statistical analysis methods: the first correlation degree ranking column vector obtained by the correlation coefficient matrix method, the second correlation degree ranking column vector obtained by the distance correlation coefficient, the third correlation degree ranking column vector obtained by the recursive feature elimination method, and the fourth correlation degree ranking column vector obtained by the grey relational analysis method. Each statistical analysis method has its own advantages and limitations. Therefore, by weighted combination, the deficiencies of a single method can be overcome, and a more robust feature evaluation result can be provided. The specifically constructed feature comprehensive screening model is Among them, β idenotes the weight assigned to the i-th statistical analysis method; y i denotes the column vector of relevance rankings obtained by the i-th statistical analysis method; the weight β here i reflects the relative importance of each statistical analysis method in the overall evaluation and can be set according to the specific application scenario and the performance of the method. For example, the correlation coefficient matrix method and the distance correlation coefficient perform excellently in dealing with linear and non-linear relationships. Therefore, the weights of the correlation coefficient matrix method and the distance correlation coefficient may be set relatively high. The recursive feature elimination method and the grey relational analysis method are particularly effective in dealing with highly correlated features and time series data. Therefore, the weights of the recursive feature elimination method and the grey relational analysis method can be set relatively low, but still contribute to the final result.
[0083] By solving the feature comprehensive screening model, a comprehensive relevance column vector is obtained. The comprehensive relevance column vector contains the comprehensive relevance evaluation values of all feature information, arranged in descending order, revealing the ranking of the features' impact on commission traceability. Features with high relevance rank higher in the column vector, indicating that they have higher importance in the automated commission traceability check, while features with low relevance rank lower, indicating that they contribute less to commission traceability.
[0084] The finally obtained comprehensive relevance column vector can be used for feature selection, that is, selecting the top-ranked features for subsequent model training and commission traceability check. These features with high relevance are considered to have a significant impact on the accuracy of commission traceability and are the key to building a prediction model.
[0085] The construction and solution process of the feature comprehensive screening model not only improves the accuracy of feature selection but also ensures the reliability and efficiency of the automated commission traceability check method. By integrating the evaluation results of different methods, the bias that may be brought by a single method can be avoided, and a more comprehensive and stable set of important features can be obtained.
[0086] In summary, by weighted combination of the relevance ranking vectors obtained from multiple statistical analysis methods and solving to obtain the comprehensive relevance column vector, the automated commission traceability check method can more scientifically and reasonably select features, improve the prediction performance of the model and the automation level of commission traceability, and contribute to more accurate data analysis and decision-making in commission traceability.
[0087] In another embodiment of the present application, a machine learning algorithm based on gradient boosting decision trees is used to construct an initial commission traceability prediction model. The above-mentioned target feature information is input into the above-mentioned initial commission traceability prediction model for training and optimization to obtain a target commission traceability prediction model, including: a first construction step of using the LightGBM algorithm to construct the above-mentioned initial commission traceability prediction model; a first calculation step of inputting the above-mentioned target feature information into the above-mentioned initial commission traceability prediction model for training and calculating the loss function of the above-mentioned initial commission traceability prediction model, where the loss function is used to quantify the difference between the predicted value and the true value of the above-mentioned initial commission traceability prediction model; a second calculation step of calculating, for each sample data point, the negative gradient of the above-mentioned loss function with respect to the above-mentioned predicted value, where the negative gradient represents the optimization direction of the above-mentioned initial commission traceability prediction model at the above-mentioned sample data point; a second construction step of constructing a new decision tree according to the above-mentioned negative gradient, where the goal of the above-mentioned new decision tree is to fit the above-mentioned negative gradient to reduce the above-mentioned loss function; an update step of adding the above-mentioned new decision tree to the above-mentioned initial commission traceability prediction model and updating the predicted value of the above-mentioned initial commission traceability prediction model; repeating the above-mentioned first calculation step, the above-mentioned second calculation step, the above-mentioned second construction step, and the above-mentioned update step until a preset stop condition is reached to obtain the above-mentioned target commission traceability prediction model.
[0088] Specifically, the embodiment of the present application uses the LightGBM algorithm to construct an initial commission traceability prediction model. LightGBM is an efficient distributed framework based on gradient boosting (GBDT), and its main features are efficient model training and prediction capabilities, suitable for processing large-scale data sets. The LightGBM algorithm gradually enhances the prediction ability of the model by constructing a series of weak learners (decision trees). During the construction of each tree, the model updates the weights of the samples according to the current error of the model. LightGBM reduces the error step by step through multiple rounds of iteration, fitting each new tree to the residuals of the previous model.
[0089] When constructing the initial commission traceability prediction model, the LightGBM framework is used to train the first decision tree based on the historical data of commission traceability and the target feature information. This decision tree will learn how to predict the accuracy of commission traceability from the feature information. The target feature information is input into the constructed initial commission traceability prediction model, and the model will make predictions on these features to generate predicted values. Then, the loss function is calculated. The loss function is a measure that quantifies the difference between the predicted value and the true value of the model, and it reflects the accuracy of the model prediction. In LightGBM, the choice of the loss function depends on the specific prediction problem. For regression problems, it can be the squared loss function, and for classification problems, it can be the cross-entropy loss function.
[0090] For each sample data point, it is necessary to calculate the negative gradient of the loss function with respect to the model's predicted value. The negative gradient gives the optimization direction of the model on this sample, that is, how to adjust the model's prediction to reduce the loss function. Simply put, the negative gradient is the derivative of the difference between the model's predicted value and the true value, indicating the direction of adjustment of the model parameters.
[0091] Based on the obtained negative gradient, a new decision tree is constructed. The goal of the new tree is to fit these negative gradients, that is, the new tree will be trained to predict the negative gradient of each sample, and the ultimate goal is to reduce the loss function and improve the prediction accuracy of the model.
[0092] The constructed new decision tree is added to the existing model and combined with the previous decision trees to update the prediction ability of the model. This process is equivalent to the update of the model parameters. Each new decision tree is added to correct the deficiencies in the prediction of the existing model. The predicted value is also updated to reflect the contribution of the newly added decision tree.
[0093] The above steps are repeated until a preset stopping condition is met. The stopping condition can be reaching a predetermined number of iterations or the loss function no longer decreasing significantly. In each iteration, model training, negative gradient calculation, new tree construction, and model update are repeated. This process is called gradient boosting.
[0094] Finally, after multiple rounds of iteration and optimization, a target commission retrospective prediction model is obtained. This model has been fully trained and improved in terms of the accuracy of predicting commission retrospective, can effectively predict new commission data, and at the same time, the generalization ability of the model has been enhanced, and it can better handle unseen data.
[0095] The whole process utilizes the efficiency and flexibility of the LightGBM algorithm. In view of the characteristics of commission retrospective, the model is continuously optimized to improve the prediction accuracy, thus providing powerful data processing and decision-making support capabilities for automated commission retrospective inspection.
[0096] In a specific embodiment, a machine learning algorithm based on gradient boosting decision tree is used to construct an initial commission retrospective prediction model, and the above target feature information is input into the above initial commission retrospective prediction model for training and optimization to obtain a target commission retrospective prediction model. It also includes: using the Leaf-wise decision tree growth strategy and the histogram-based gradient boosting method to optimize the above target commission retrospective prediction model.
[0097] Specifically, the traditional gradient boosting decision tree algorithm adopts the Level-wise strategy when constructing a decision tree, that is, constructing the decision tree layer by layer. Starting from the root node of the tree, it grows layer by layer until the preset depth or stopping condition is reached. For the schematic diagram of the Level-wise tree growth strategy, see Figure 3However, this method leads to some unnecessary computations because all leaf nodes are grown simultaneously, even though some leaf nodes contribute little to reducing the loss function. This strategy introduces a large amount of computational overhead in practical applications, especially when the eigenvalue distribution is uneven, and nodes in some layers may not need to be split to achieve good segmentation results.
[0098] In contrast, the Leaf-wise decision tree growth strategy (also known as the best-first decision tree growth strategy) always starts growing from the best split point of the current leaf node. This means that it preferentially selects the leaf nodes that contribute the most to reducing the loss function for growth, ensuring that each tree is constructed based on the leaf nodes that maximize the improvement of the model performance, thereby significantly reducing unnecessary computations. See the schematic diagram of the Leaf-wise tree growth strategy in Figure 4 The benefit of this is that it can significantly reduce ineffective splits, improve the model construction speed, and at the same tree depth, the Leaf-wise strategy often can construct a better (i.e., with less loss) model. To avoid overfitting caused by overgrowth, LightGBM usually sets some limiting conditions, such as restricting the maximum depth of the tree, the minimum number of samples in a leaf node, etc.
[0099] The histogram-based gradient boosting method is one of the core features of the LightGBM algorithm, which significantly improves the model training speed while reducing the memory consumption. In traditional GBDT, the splitting of decision tree nodes is done by finding an optimal split point. Specifically, for each eigenvalue, all possible split points are traversed, and then the decrease in the final loss function (such as squared loss or cross-entropy) is calculated to find a split point that maximizes the loss decrease. Although this method is simple, it becomes very time-consuming and computationally resource-intensive when dealing with a large number of features and data because it requires scanning and calculating every possible split value for each feature.
[0100] To reduce the computational complexity of finding split points during the decision tree construction, LightGBM proposes an equalization splitting method based on histograms. The core idea of this method is to discretize the continuous eigenvalue and divide it into a finite number of discrete bins, construct a histogram for each feature, and find an approximate optimal split point based on the histogram. See the schematic diagram of finding the optimal split point in Figure 5 This greatly reduces the computational time for split point search and also avoids excessive dependence on the precision of floating-point features. The specific steps are as follows:
[0101] Discretized Features: First, for each continuous feature value, LightGBM discretizes the values according to certain rules (such as bucketing or quantization). Usually, LightGBM divides the range of the feature values into equal-frequency bins, and each bin represents a group of continuous feature values within a specific range.
[0102] Histogram Construction: For the discretized features, count the number of data samples, gradients, and weights in each bin. That is, after mapping all samples to specific bins, calculate the sum of gradients and second-order derivatives for each bin.
[0103] Finding the Optimal Split Point: Using the statistical data in these bins, scan different bins in sequence, and find the split point that maximizes the decrease in the loss function by calculating the gradient descent effect when each bin is used as a split point.
[0104] Decision Tree Splitting: According to the found optimal split point, split the current leaf node into two child nodes, and continue to apply the Leaf-wise strategy to select the next optimal split node.
[0105] Through the Leaf-wise decision tree growth strategy and the histogram-based gradient boosting method, LightGBM can quickly and efficiently construct and optimize the commission traceability prediction model, and maintain a high training speed and prediction accuracy even when dealing with large-scale datasets. The use of these optimization strategies not only accelerates the model training process but also improves the generalization ability of the model, enabling the final commission traceability prediction model to more accurately predict the correctness of commission operations and compliance during the traceability period.
[0106] In summary, combining the Leaf-wise decision tree growth strategy and the histogram-based gradient boosting method can significantly improve the training efficiency and prediction accuracy of the automated commission traceability prediction model, providing strong technical support for the automation of commission traceability.
[0107] In another embodiment of the present application, after obtaining the target commission traceability prediction model, the above method further includes: using the coefficient of determination root mean square error mean absolute error and mean absolute percentage error to evaluate the accuracy of the above target commission traceability prediction model, where n represents the number of samples, y i represents the i-th measured value, represents the i-th predicted value, represents the mean of the above measured values.
[0108] Specifically, the coefficient of determination R 2 is used to measure the ability of the model to explain data changes. R2 The value is between 0 and 1. The closer the value is to 1, the stronger the model's ability to explain data changes and the better the prediction effect. By calculating the coefficient of determination R 2 , the goodness of fit between the model's predicted values and the true values can be evaluated. The root mean square error RMSE is a method to measure the magnitude of the error between the predicted values and the true values, especially suitable for numerical prediction tasks. The smaller the RMSE, the better the prediction performance of the model. The mean absolute error MAE directly measures the average absolute difference between the predicted values and the true values, is more sensitive to the magnitude of the prediction error, and is not affected by positive or negative errors. The smaller the MAE value, the smaller the error of the model during prediction and the closer the prediction result is to the true value. The mean absolute percentage error MAPE is used to measure the average of the relative errors between the predicted values and the true values, and is often used to evaluate the accuracy of ratio or percentage predictions. The smaller the MAPE value, the higher the accuracy of the model prediction.
[0109] These metrics work together to evaluate the model's prediction ability from different perspectives. For example, an R 2 value close to 1 indicates that the model explains most of the data changes. However, in actual prediction, it is also necessary to pay attention to metrics such as RMSE, MAE, and MAPE that more directly reflect the magnitude of the prediction error to ensure that the model not only performs well in fitting the data but also makes accurate predictions on new data.
[0110] Specifically in the commission traceback prediction model, by calculating these evaluation metrics, the accuracy of the model in predicting commission traceback operations can be evaluated. For example, if the MAE and RMSE values are small, the MAPE value is also at a low level, and the R 2 is close to 1, this indicates that the model can well predict the correctness and compliance of commission traceback, which helps to improve the efficiency and accuracy of automated commission traceback checks.
[0111] Finally, using these evaluation metrics, the model can be evaluated regularly, and the algorithm parameters can be adjusted or the model structure can be optimized according to the model's performance to continuously improve the model's performance in commission traceback prediction.
[0112] In the embodiments of this application, after the model training is completed, new commission data is input into the trained model. The model will predict the input data, mark the suspected incorrect commission records, and finally output the marked results for manual review and further processing.
[0113] To improve the reliability and accuracy of the model, it is also necessary to manually verify the results output by the model. Manually review the suspected incorrect records marked by the model to confirm whether there are indeed errors. And feedback the results of the manual review to the model to adjust the model parameters and improve the model performance. Based on the feedback results, continuously optimize the model to improve its performance in actual applications.
[0114] The present application also provides a specific embodiment of the automated commission traceability inspection method, which specifically includes the following steps:
[0115] Step 1: Data collection. Collect and organize the settlement data of historical accounting periods and the adjustment data of the current accounting period in the partner settlement subsystem.
[0116] Step 2: Data preprocessing. Clean and organize the commission traceability data, including data conversion and data deduplication.
[0117] Step 3: Data splitting. Split the data into a training set, a validation set, and a test set according to the ratio of 8:1:1, that is, split the data set into a training set, a validation set, and a test set, where the training set accounts for 80%, and the validation set and the test set each account for 10%, to ensure that the model can better learn the data features during training, and at the same time evaluate and optimize the model through the validation set and the test set.
[0118] Step 4: Feature extraction and optimization strategy. Apply a feature comprehensive screening model to calculate the information values of a large number of feature value fields, and take the top 20 feature fields with the highest information values as the feature values for subsequent model training.
[0119] Step 5: Model training and optimization. Use the training set to train the LightGBM model, and optimize the model parameters through multiple iterations to maximize the prediction accuracy. During parameter tuning and monitoring the training process, adjust the hyperparameters of the model through methods such as grid search or random search to obtain the best model performance.
[0120] Step 6: Model evaluation metrics. Real-time monitor the loss function and evaluation metrics during the training process to ensure that the model is gradually optimized.
[0121] Step 7: Model verification. Use the test set to verify the trained model and evaluate the prediction performance and generalization ability of the model.
[0122] Step 8: Result output. Output the automated comparison results of commission traceability, including the traceability accuracy rate, etc., and mark the suspected incorrect commission records.
[0123] The embodiment of the present application also provides an automated commission traceability inspection device. It should be noted that the automated commission traceability inspection device of the embodiment of the present application can be used to execute the automated commission traceability inspection method provided by the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0124] The following introduces the automated commission traceability inspection device provided by the embodiments of the present application.
[0125] Figure 6 It is a structural block diagram of the automated commission traceability inspection device according to the embodiments of the present application. As Figure 6 shown, the device includes an acquisition unit 10, a determination unit 20, a training and optimization unit 30, and an output unit 40.
[0126] The acquisition unit 10 is configured to obtain historical commission data from a data source and preprocess the historical commission data to obtain preprocessed historical commission data, where the historical commission data includes various feature information related to commission traceability;
[0127] The determination unit 20 is configured to perform feature analysis on the preprocessed historical commission data by using various statistical analysis methods to obtain a comprehensive relevance ranking of various pieces of the feature information, and determine the feature information within the first preset range of the comprehensive relevance ranking as target feature information;
[0128] The training and optimization unit 30 is configured to construct an initial commission traceability prediction model by using a machine learning algorithm based on gradient boosting decision trees, input the target feature information into the initial commission traceability prediction model for training and optimization, and obtain a target commission traceability prediction model;
[0129] The output unit 40 is configured to input current commission data into the target commission traceability prediction model for prediction to obtain a prediction result, mark the data records whose prediction results are not within the second preset range as abnormal data records, and output the abnormal data records.
[0130] In an embodiment of the present application, the determination unit includes:
[0131] A feature analysis module, configured to perform feature analysis on the preprocessed historical commission data by using various statistical analysis methods to obtain multiple relevance degrees of each piece of the feature information with commission traceability, and arrange the values of the multiple relevance degrees obtained by each statistical analysis method from large to small to obtain a relevance ranking column vector corresponding to each statistical analysis method;
[0132] A construction module, configured to perform weighted construction on the relevance ranking column vectors corresponding to each statistical analysis method to construct a feature comprehensive screening model, solve the feature comprehensive screening model, and obtain a comprehensive relevance column vector of the feature information;
[0133] A determination module, configured to determine the feature information within the first preset range of the comprehensive relevance column vector as target feature information.
[0134] In a specific embodiment, the above-mentioned feature analysis module includes:
[0135] The first calculation sub-module is used to calculate multiple first correlation degrees between each of the above-mentioned feature information and the commission traceability variable by using the correlation coefficient matrix method, and arrange the values of the multiple first correlation degrees from large to small to obtain a first correlation degree ranking column vector;
[0136] The second calculation sub-module is used to calculate a second correlation degree between each of the above-mentioned feature information and the above-mentioned commission traceability variable by using the distance correlation coefficient, and arrange the values of the multiple second correlation degrees from large to small to obtain a second correlation degree ranking column vector;
[0137] The third calculation sub-module is used to calculate a third correlation degree between each of the above-mentioned feature information and the above-mentioned commission traceability variable by using the recursive feature elimination method, and arrange the values of the multiple third correlation degrees from large to small to obtain a third correlation degree ranking column vector;
[0138] The fourth calculation sub-module is used to calculate a fourth correlation degree between each of the above-mentioned feature information and the above-mentioned commission traceability variable by using the grey relational analysis method, and arrange the values of the multiple fourth correlation degrees from large to small to obtain a fourth correlation degree ranking column vector.
[0139] In another specific embodiment, the above-mentioned construction module includes:
[0140] The construction sub-module is used to construct a feature comprehensive screening model by weighting the above-mentioned correlation degree ranking column vectors where Y is the above-mentioned comprehensive correlation degree column vector, and β i represents the weight assigned to the i-th above-mentioned statistical analysis method; y i represents the correlation degree ranking column vector obtained by the i-th above-mentioned statistical analysis method;
[0141] The solution sub-module is used to solve the feature comprehensive screening model to obtain the above-mentioned comprehensive correlation degree column vector arranged in descending order of correlation degree.
[0142] In another embodiment of the present application, the above-mentioned training and optimization unit includes:
[0143] The first construction module is used to construct the above-mentioned initial commission traceability prediction model by using the LightGBM algorithm;
[0144] The first calculation module is used to input the above-mentioned target feature information into the above-mentioned initial commission traceability prediction model for training, and calculate the loss function of the above-mentioned initial commission traceability prediction model, and the loss function is used to quantify the difference between the predicted value and the true value of the above-mentioned initial commission traceability prediction model;
[0145] A second calculation module, configured to calculate, for each sample data point, the negative gradient of the above loss function with respect to the above predicted value, where the negative gradient represents the optimization direction of the above initial commission retrospective prediction model at the above sample data point;
[0146] A second construction module, configured to construct a new decision tree according to the above negative gradient, where the goal of the above new decision tree is to fit the above negative gradient to reduce the above loss function;
[0147] An update module, configured to add the above new decision tree to the above initial commission retrospective prediction model and update the predicted value of the above initial commission retrospective prediction model;
[0148] Repeat the execution of the above first calculation module, the above second calculation module, the above second construction module, and the above update module until a preset stop condition is reached, and obtain the above target commission retrospective prediction model.
[0149] In a specific embodiment, the above training and optimization unit further includes:
[0150] An optimization module, configured to optimize the above target commission retrospective prediction model by using a Leaf-wise decision tree growth strategy and a histogram-based gradient boosting method.
[0151] In another embodiment of the present application, the above device further includes:
[0152] An evaluation unit, configured to adopt the coefficient of determination root mean square error mean absolute error and mean absolute percentage error to evaluate the accuracy of the above target commission retrospective prediction model, where n represents the number of samples, y i represents the i-th measured value, represents the i-th predicted value, represents the mean of the above measured values.
[0153] The above automatic commission retrospective inspection device includes a processor and a memory. The above acquisition unit, determination unit, training and optimization unit, output unit, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions. The above modules are all located in the same processor; or, the above respective modules are located in different processors in any combination form.
[0154] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0155] An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned automated commission traceability inspection method.
[0156] Specifically, the automated commission traceability inspection method includes:
[0157] Step S201: Obtain historical commission data from a data source, and preprocess the historical commission data to obtain preprocessed historical commission data. The historical commission data includes various feature information related to commission traceability.
[0158] Step S202: Use various statistical analysis methods to perform feature analysis on the preprocessed historical commission data, obtain a comprehensive relevance ranking of various pieces of the feature information, and determine the feature information with the comprehensive relevance ranking within a first preset range as the target feature information.
[0159] Step S203: Use a machine learning algorithm based on gradient boosting decision trees to construct an initial commission traceability prediction model, input the target feature information into the initial commission traceability prediction model for training and optimization to obtain a target commission traceability prediction model.
[0160] Step S204: Input the current commission data into the target commission traceability prediction model for prediction to obtain a prediction result, mark the data records whose prediction results are not within a second preset range as abnormal data records, and output the abnormal data records.
[0161] An embodiment of the present invention provides an electronic device, including a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the following steps:
[0162] Step S201: Obtain historical commission data from a data source, and preprocess the historical commission data to obtain preprocessed historical commission data. The historical commission data includes various feature information related to commission traceability.
[0163] Step S202: Use various statistical analysis methods to perform feature analysis on the preprocessed historical commission data, obtain a comprehensive relevance ranking of various pieces of the feature information, and determine the feature information with the comprehensive relevance ranking within a first preset range as the target feature information.
[0164] Step S203: Use a machine learning algorithm based on gradient boosting decision trees to construct an initial commission traceability prediction model. Input the above-mentioned target feature information into the initial commission traceability prediction model for training and optimization to obtain a target commission traceability prediction model.
[0165] Step S204: Input the current commission data into the above-mentioned target commission traceability prediction model for prediction to obtain a prediction result. Mark the data records whose prediction results are not within the second preset range as abnormal data records, and output the above-mentioned abnormal data records.
[0166] This application also provides a computer program product, which is suitable for executing a program initialized with at least the following method steps when executed on a data processing device:
[0167] Step S201: Obtain historical commission data from a data source and preprocess the historical commission data to obtain preprocessed historical commission data. Among them, the historical commission data includes various feature information related to commission traceability.
[0168] Step S202: Use various statistical analysis methods to perform feature analysis on the preprocessed historical commission data to obtain a comprehensive correlation ranking of various feature information, and determine the feature information whose comprehensive correlation ranking is within the first preset range as target feature information.
[0169] Step S203: Use a machine learning algorithm based on gradient boosting decision trees to construct an initial commission traceability prediction model. Input the above-mentioned target feature information into the initial commission traceability prediction model for training and optimization to obtain a target commission traceability prediction model.
[0170] Step S204: Input the current commission data into the above-mentioned target commission traceability prediction model for prediction to obtain a prediction result. Mark the data records whose prediction results are not within the second preset range as abnormal data records, and output the above-mentioned abnormal data records.
[0171] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0172] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0173] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0174] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0176] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0177] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0178] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0179] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0180] The foregoing is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. An automated commission retrospective checking method, characterized in that: include: Acquire historical commission data from a data source, and pre-process the historical commission data to obtain pre-processed historical commission data, wherein the historical commission data includes a variety of feature information related to commission tracing; perform feature analysis on the pre-processed historical commission data using a variety of statistical analysis methods to obtain comprehensive relevance rankings of the various feature information, and determine the feature information with the comprehensive relevance ranking within a first preset range as target feature information; Using a machine learning algorithm based on a gradient boosting decision tree to build an initial commission tracing prediction model, inputting the target feature information into the initial commission tracing prediction model for training and optimization, and obtaining a target commission tracing prediction model; The current commission data is input into the target commission retrospective prediction model for prediction to obtain a prediction result, and the data record whose prediction result is not within the second preset range is marked as an abnormal data record, and the abnormal data record is output.
2. The method according to claim 1, characterized in that: Using a variety of statistical analysis methods to perform feature analysis on the preprocessed historical commission data to obtain comprehensive relevance rankings of a variety of feature information, and determining the feature information within a first preset range of the comprehensive relevance ranking as target feature information, including: Using multiple statistical analysis methods to perform feature analysis on the preprocessed historical commission data, obtaining multiple correlations between each feature information and commission tracing, and arranging the multiple correlation values obtained by each statistical analysis method from large to small to obtain a correlation ranking column vector corresponding to each statistical analysis method; The correlation ranking column vector corresponding to each statistical analysis method is weighted to construct a feature comprehensive screening model, and the feature comprehensive screening model is solved to obtain a comprehensive correlation column vector of the feature information; The feature information in the comprehensive correlation column vector within the first preset range is determined as target feature information.
3. The method according to claim 2, characterized in that A plurality of statistical analysis methods are used to perform feature analysis on the preprocessed historical commission data to obtain a plurality of correlations between each feature information and commission tracing, and the values of the plurality of correlations obtained by each statistical analysis method are arranged from large to small to obtain a correlation ranking column vector corresponding to each statistical analysis method, including: A correlation coefficient matrix method is used to calculate a plurality of first correlations between each of the feature information and the commission tracing variable, and the values of the plurality of first correlations are arranged from large to small to obtain a first correlation ranking column vector; The second correlation between each of the feature information and the commission tracing variable is calculated by using a distance correlation coefficient, and the values of the plurality of the second correlations are arranged from large to small to obtain a second correlation ranking column vector; A recursive feature elimination method is used to calculate the third correlation between each feature information and the commission tracing variable, and multiple third correlation values are arranged from large to small to obtain a third correlation ranking column vector; The fourth correlation between each of the feature information and the commission tracing variable is calculated using a grey correlation analysis method, and multiple fourth correlation values are arranged from large to small to obtain a fourth correlation ranking column vector.
4. The method according to claim 2, characterized in that: The correlation ranking column vector corresponding to each statistical analysis method is weighted to construct a feature comprehensive screening model, and the feature comprehensive screening model is solved to obtain a comprehensive correlation column vector of the feature information, including: By weighting the correlation ranking column vector, a feature comprehensive screening model is constructed. Wherein, Y is the comprehensive correlation column vector, β i represents the weight assigned to the i-th statistical analysis method; y i represents the correlation ranking column vector obtained by the i-th statistical analysis method; By solving the feature comprehensive screening model, the comprehensive correlation column vector arranged from high to low correlation is obtained.
5. The method according to claim 1, characterized in that: An initial commission tracing prediction model is constructed using a machine learning algorithm based on a gradient boosting decision tree, and the target feature information is input into the initial commission tracing prediction model for training and optimization to obtain a target commission tracing prediction model, including: The first construction step is to use the LightGBM algorithm to construct the initial commission retrospective prediction model; The first calculation step is to input the target feature information into the initial commission retrospective prediction model for training, and calculate the loss function of the initial commission retrospective prediction model, wherein the loss function is used to quantify the difference between the predicted value and the true value of the initial commission retrospective prediction model; A second calculation step, for each sample data point, calculating a negative gradient of the loss function with respect to the predicted value, wherein the negative gradient represents an optimization direction of the initial commission retrospective prediction model on the sample data point; A second construction step is to construct a new decision tree according to the negative gradient, wherein the goal of the new decision tree is to fit the negative gradient to reduce the loss function; An updating step, adding the new decision tree to the initial commission retrospective prediction model, and updating the prediction value of the initial commission retrospective prediction model; The first calculation step, the second calculation step, the second construction step and the updating step are repeatedly performed until a preset stop condition is reached to obtain the target commission retrospective prediction model.
6. The method according to claim 5, characterized in that An initial commission tracing prediction model is constructed using a machine learning algorithm based on a gradient boosting decision tree, and the target feature information is input into the initial commission tracing prediction model for training and optimization to obtain a target commission tracing prediction model, further comprising: The target commission retrospective prediction model is optimized using a Leaf-wise decision tree growth strategy and a histogram-based gradient boosting method.
7. The method according to claim 1, characterized in that After obtaining the target commission retrospective prediction model, the method further includes: Coefficient of determination Root mean square error Mean absolute error and mean absolute percentage error Evaluate the accuracy of the target commission retrospective prediction model, where n represents the number of samples, y i represents the i-th measured value, represents the i-th predicted value, represents the mean of the measured values.
8. An automated commission retrospective inspection device, characterized in that: include: An acquisition unit, configured to acquire historical commission data from a data source, and preprocess the historical commission data to obtain preprocessed historical commission data, wherein the historical commission data includes a variety of feature information related to commission tracing; a determination unit, configured to perform feature analysis on the preprocessed historical commission data using a plurality of statistical analysis methods, obtain comprehensive relevance rankings of a plurality of feature information, and determine the feature information within a first preset range of the comprehensive relevance rankings as target feature information; A training and optimization unit, configured to construct an initial commission tracing prediction model using a machine learning algorithm based on a gradient boosting decision tree, input the target feature information into the initial commission tracing prediction model for training and optimization, and obtain a target commission tracing prediction model; The output unit is used to input the current commission data into the target commission retrospective prediction model for prediction, obtain the prediction result, mark the data record whose prediction result is not within the second preset range as an abnormal data record, and output the abnormal data record.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the automated commission retrospective checking method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing the automated commission retrospective checking method according to any one of claims 1 to 7.