Business data processing method and device, electronic equipment, computer readable storage medium and computer program product
By building a decision tree to identify the influencing factors of business parameter changes, the problem of insufficient flexibility in existing technologies is solved, flexible adaptation and accurate analysis in different business scenarios and data characteristics are achieved, and the versatility of data processing and risk monitoring capabilities are improved.
Patent Information
- Application Number
- CN202510605751.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies are difficult to adapt flexibly to different business scenarios and data characteristics, and are unable to accurately capture the influencing factors of changes in business parameters, resulting in insufficient versatility and flexibility in data processing methods.
By constructing a decision tree corresponding to the business parameters, the decision path before and after the business parameters change is determined. The influencing factors and decision trees are used to automatically capture nonlinear relationships, identify the third influencing factor, and analyze the changes in business parameters.
It improves the flexibility and versatility of data processing methods in actual business applications, can accurately identify influencing factors when business parameters change, and supports risk monitoring and strategy adjustments.
Smart Images

Figure CN120687947A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device, computer-readable storage medium, and computer program product for processing business data. Background Art
[0002] In business data analysis across various industries, changes in business operating status are often the result of multiple factors, including business operations, market environment, technical conditions, and management decisions. These factors often involve business parameters, meaning that changes in these parameters often reflect the business's operating status. When significant changes in business parameters occur, in-depth analysis and explanation of the causes of these changes are often necessary to promptly identify issues, optimize decisions, and implement targeted measures. This analytical process is crucial for operational management, risk control, strategic planning, and policy formulation. Summary of the Invention
[0003] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium, and computer program product for processing business data, which can be applied to different business scenarios and different data characteristics, thereby improving the flexibility and versatility of the data processing method in actual business applications.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] This embodiment of the present application provides a method for processing business data, the method comprising:
[0006] Determine a business parameter used to reflect the operating status of a first business, and a first decision tree corresponding to the business parameter, wherein the first decision tree includes multiple decision paths; when the parameter value of the business parameter changes from a first parameter value to a second parameter value, based on the first decision path matching the first parameter value in the first decision tree, determine the first impact value of the first influencing factor in the first decision path on the business parameter, and based on the second decision path matching the second parameter value in the first decision tree, determine the second impact value of the second influencing factor in the second decision path on the business parameter; based on the first impact value and the second impact value, determine a third impact factor for the change in the parameter value of the business parameter from the first impact factor and the second impact factor.
[0007] This embodiment of the present application provides a device for processing business data, including:
[0008] A first determining module, configured to determine a service parameter reflecting an operating state of a first service and a first decision tree corresponding to the service parameter, wherein the first decision tree includes a plurality of decision paths;
[0009] a second determining module, configured to determine, when the parameter value of the service parameter changes from the first parameter value to the second parameter value, based on a first decision path in the first decision tree that matches the first parameter value, a first impact value of a first influencing factor in the first decision path on the service parameter, and determine, based on a second decision path in the first decision tree that matches the second parameter value, a second impact value of a second influencing factor in the second decision path on the service parameter;
[0010] The third determining module is configured to determine, based on the first impact value and the second impact value, a third impact factor causing a change in the parameter value of the service parameter from among the first impact factor and the second impact factor.
[0011] An embodiment of the present application provides an electronic device, comprising:
[0012] a memory for storing computer-executable instructions or computer programs;
[0013] The processor is used to implement the business data processing method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0014] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing a method for processing business data provided in an embodiment of the present application when executed by a processor.
[0015] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the business data processing method provided in the embodiment of the present application is implemented.
[0016] The embodiment of the present application has the following beneficial effects: through the above method, when performing data analysis on the first business, the business parameters reflecting the operating status of the first business are determined, and the first decision tree corresponding to the business parameters is determined. When the parameter value of the business parameter changes from the first parameter value to the second parameter value, a first decision path matching the first parameter value and a second decision path matching the second parameter value can be determined in the first decision tree, thereby determining the first impact value of the first influencing factor in the first decision path on the business parameter and the second impact value of the second influencing factor in the second decision path on the business parameter. Finally, based on the first impact value and the second impact value, the first impact factor and the second impact factor can be used to determine the impact value of the business parameter. The third influencing factor of the parameter value change of the parameter, this method can use the decision tree to automatically capture the nonlinear relationship and interaction between the business parameters and the influencing factors. In this way, when analyzing the changes in the business parameters of the first business, there is no need to preset the precise mathematical expression between the business parameters and the influencing factors as in the related technology. When the business parameters change, the third influencing factor that causes the business parameters to change can be accurately captured based on the first decision tree, so that the changes in the business parameters can be analyzed according to the third influencing factor. This method of processing business data can be applicable to different business scenarios and different data characteristics, thereby improving the flexibility and versatility of the data processing method in actual business applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a schematic diagram of the architecture of the business data processing system provided in an embodiment of the present application;
[0018] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0019] Figure 3 This is a flowchart of a method for processing business data provided by an embodiment of the present application;
[0020] Figure 4 This is a flow chart of the training method of the gradient boosting tree model provided in the embodiment of the present application.
[0021] Figure 5 is a schematic diagram of a first decision tree provided in an embodiment of the present application;
[0022] Figure 6A This is a first schematic diagram of a method for determining the fifth impact factor provided in an embodiment of the present application;
[0023] Figure 6B This is a second schematic diagram of a method for determining the fifth impact factor provided in an embodiment of the present application;
[0024] Figure 7: is a corresponding schematic diagram of the key indicators and the fourth impact factor provided in the embodiment of the present application;
[0025] Figure 8 It is a schematic diagram of a first decision tree of business parameters and a fourth influencing factor provided in an embodiment of the present application.
[0026] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0028] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0029] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0030] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0031] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0032] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0033] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0034] 1) In response to: used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0035] 2) Business Parameters: These are core quantitative parameters that reflect the operational status of a business (such as performance or system operation) and require focus and evaluation in business operational status analysis across various industries. Business parameters are typically key data points measuring a specific aspect of performance, and changes in their values can directly or indirectly reflect issues or trends in the business, technology, or management.
[0036] Business parameters are closely related to specific business scenarios or technical fields, and can reflect key issues in the actual business operation process. The numerical values corresponding to business parameters (i.e., parameter values) may fluctuate with changes in time, environment or conditions, and require continuous monitoring and analysis.
[0037] Business parameters may vary across different industries. Here are some examples from application scenarios: In the retail industry's e-commerce and sales businesses, business parameters may include sales volume, customer conversion rate, and inventory turnover rate; in the financial industry's investment and risk identification businesses, business parameters may include default rate, return on investment, and customer churn rate; in the manufacturing industry's production and equipment maintenance businesses, business parameters may include production cost, equipment failure rate, and production efficiency; in the medical industry's clinical and medical services and medical quality maintenance businesses, business parameters may include patient cure rate, length of hospital stay, and medical expenses; in the internet industry's product operation and marketing businesses, business parameters may include user activity, click-through rate, and conversion rate; in the energy industry's energy audit and energy maintenance businesses, business parameters may include energy consumption, carbon emissions, and equipment utilization.
[0038] 3) Influencing factors, also known as factor indicators or independent variables, are potential factors or variables that may influence the numerical changes of business parameters (dependent variables). These are variables used during the analysis process to explain, predict, or evaluate changes in business parameters. They can be single variables or combinations of multiple variables that collectively affect the business parameters. These influencing factors can be internal or external, controllable or uncontrollable, and they act directly or indirectly on the business parameters, causing their numerical changes.
[0039] Influencing factors can cover various types of data, including numerical data (such as temperature and price), categorical data (such as region and product type), and time series data (such as daily trading volume). There may be interactions or nonlinear relationships between influencing factors, which jointly affect business parameters. Moreover, the effects of influencing factors may change with time, environment, or conditions, requiring dynamic analysis.
[0040] The influencing factors may vary in different industries. The following are some examples of application scenarios: In the e-commerce business and sales business of the retail industry, influencing factors can be the intensity of promotional activities, holiday factors, competitor prices, customer information (such as age, gender), etc.; in the investment business and risk identification business of the financial industry, influencing factors can be interest rate changes, economic policy adjustments, customer credit scores, market volatility, etc.; in the production business and equipment maintenance business of the manufacturing industry, influencing factors can be equipment service life, maintenance frequency, raw material quality, ambient temperature, etc.; in the clinical and medical service business and medical quality maintenance business of the medical industry, influencing factors can be patient age, treatment plan, medical equipment type, doctor experience level, etc.; in the product operation business and marketing promotion business of the Internet industry, influencing factors can be advertising volume, user behavior data, website loading speed, market competition, etc.
[0041] 4) Coefficient of Determination (usually denoted as R) 2 ) is a key metric used to evaluate the performance of regression tree models. It measures the model's ability to explain changes in the dependent variable (such as business parameters), that is, the extent to which the model can capture the variation in the dependent variable. The coefficient of determination ranges from [0 to 1]. Values closer to 1 indicate a better model fit, while values closer to 0 indicate a poorer model fit.
[0042] The embodiments of the present application provide a method, apparatus, device, computer-readable storage medium, and computer program product for processing business data, which can be applied to different business scenarios and different data characteristics, thereby improving the flexibility and versatility of the data processing method in actual business applications.
[0043] See also Figure 1 , Figure 1 This is an architectural diagram of a business data processing system 100 provided in an embodiment of the present application. To support a business data processing application, a terminal 401 is connected to a server 200 via a network 300. The network 300 may be a wide area network or a local area network, or a combination of the two.
[0044] It should be noted that Figure 1 The business data processing system shown can be applied to risk control scenarios as a functional module of a risk control system, wherein the risk control system can be used to monitor the risk of a first business of at least one industry. For example, the risk control system can be used to monitor several business parameters that can reflect the operating status of the first business. When a change in the parameter value of one or more business parameters is detected, the business data processing method of the embodiment of the present application (i.e., by Figure 1 The process of determining the third influencing factor that causes the parameter value of the business parameter to change is performed based on the third influencing factor, and then the risk resolution strategy is executed to stabilize the parameter value of the business parameter, or adjust the parameter value of the business parameter in a positive direction, thereby ensuring the stability of the operating status of the first business. This method can improve the stability and security of the risk control system when monitoring the risk of the first business.
[0045] Among them, risk resolution strategy refers to a series of plans, measures and actions formulated and implemented to reduce or eliminate the negative impact of changes in the parameter values of business parameters after identifying and assessing risks (i.e. determining the third influencing factor); the core goal of the risk resolution strategy is to control risks within an acceptable range through effective risk identification, thereby ensuring the smooth operation of the business and the achievement of goals.
[0046] As an example, when the risk control system is used to monitor the risks of the investment business (i.e., the first business) in the financial industry, the risk control system can monitor in real time the business parameters such as the default rate, investment return rate, and customer churn rate that can reflect the operating status of the investment business. When the risk control system detects that the parameter value of the business parameter of the customer churn rate has changed, for example, the customer churn rate has increased (from 20% to 60%), the terminal 401 can generate a parameter analysis request for the business parameter of the customer churn rate to the server 200. In response to the parameter analysis request, the server 200 determines the business parameter (customer churn rate) used to reflect the operating status of the first business (investment business), and the first decision tree corresponding to the business parameter (customer churn rate), the first decision tree including multiple decision paths; when the parameter value of the business parameter (customer churn rate) changes from the first parameter value (20%) to the second parameter value (60%), based on the first decision path in the first decision tree that matches the first parameter value , determine the first impact value of the first influencing factor in the first decision path on the business parameter (customer churn rate), and based on the second decision path matching the second parameter value in the first decision tree, determine the second impact value of the second influencing factor in the second decision path on the business parameter (customer churn rate); based on the first impact value and the second impact value, determine the third influencing factor that causes the parameter value of the business parameter (customer churn rate) to change from the first influencing factor and the second influencing factor, and then determine the corresponding risk resolution strategy based on the third influencing factor, and send the risk resolution strategy to the terminal 401, so that the terminal 401 adjusts the subsequent operation of the first business based on the risk resolution strategy, so that the business parameter (customer churn rate) can be reduced.
[0047] Here, for different influencing factors of business parameters, corresponding risk resolution strategies can be pre-configured for each influencing factor. In this way, after determining the third influencing factor, the corresponding risk resolution strategy can be directly determined based on the correspondence between the pre-configured influencing factors and risk resolution strategies. In the absence of a pre-configured correspondence between influencing factors and risk resolution strategies, after determining the third influencing factor, the third influencing factor can be analyzed using the large language model to generate a risk resolution solution for the third influencing factor, and the risk resolution solution generated by the large language model can be used as the risk resolution strategy.
[0048] In some embodiments, the terminal 401 can be implemented as various types of terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart TV, a car terminal, etc., and can also be implemented as a server.
[0049] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.
[0050] See also Figure 2 , Figure 2 is a structural diagram of an electronic device 400 provided in an embodiment of the present application, Figure 2 The electronic device 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .
[0051] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0052] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0053] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0054] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0055] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0056] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0057] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB);
[0058] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0059] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.
[0060] In some embodiments, the service data processing device provided in the embodiments of the present application can be implemented in software. Figure 2 The processing device 455 for business data stored in the memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a first determination module 4551, a second determination module 4552, and a third determination module 4553. These modules are logical and can be arbitrarily combined or further separated according to the functions implemented. The functions of each module will be described below.
[0061] In other embodiments, the business data processing device provided in the embodiments of the present application can be implemented in hardware. As an example, the business data processing device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the business data processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0062] In some embodiments, the terminal or server can implement the method for processing business data provided by the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a data analysis APP; it can also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to a browser environment to run. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.
[0063] The following describes the method for processing business data provided by the embodiments of the present application with reference to the accompanying drawings. As previously mentioned, the electronic device 400 that implements the method for processing business data in the embodiments of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeatedly described below.
[0064] The method for processing business data in the embodiment of the present application is described by taking the execution subject as a server as an example. Figure 3 , Figure 3 This is a flow chart of the method for processing business data provided by the embodiment of the present application, which will be combined with Figure 3 The steps shown are explained.
[0065] In step 101, a service parameter for reflecting the operating status of a first service and a first decision tree corresponding to the service parameter are determined.
[0066] The first decision tree includes multiple decision paths.
[0067] Here, the first business may be different in different industries. For example, in the retail industry, the first business may be e-commerce business, sales business, etc.; in the financial industry, the first business may be investment business, risk identification business, etc.; in the Internet industry, the first business may be product operation business, marketing promotion business, etc.
[0068] It's understood that business parameters are core quantitative parameters that require focus and evaluation in industry analysis and reflect the operational status, performance, or system operation of a business. Business parameters are typically key data points measuring a specific aspect of performance, and changes in their values can directly or indirectly reflect issues or trends in the primary business's operations.
[0069] As an example, for different first businesses, the business parameters that can reflect the operating status of the first business are different. For example, if the first business is an e-commerce business, the business parameters that can reflect the operating status of the e-commerce business may include sales, customer conversion rate, inventory turnover rate and other parameters; if the first business is an investment business in the financial industry, the business parameters that can reflect the operating status of the investment business may include user activity, click-through rate, conversion rate and other parameters.
[0070] In actual implementation, a first decision tree of the service parameters may be pre-constructed for each service parameter of the first service, or a corresponding first decision tree may be constructed for the service parameter when a change in the value of a service parameter is detected.
[0071] In some embodiments, the first decision tree corresponding to the business parameter can be determined in the following manner: based on a plurality of fourth influencing factors preset for constructing the first decision tree, a data set is constructed, the data set including the first numerical value of each fourth influencing factor at different time points, and the third parameter value of the business parameter at different time points; based on the data set, the relationship between the business parameter and the plurality of fourth influencing factors is fitted to obtain a second decision tree; the determination coefficient of the second decision tree is determined, and the structure of the second decision tree is adjusted based on the determination coefficient to obtain a first decision tree.
[0072] Different business parameters have different influencing factors that can affect changes in their values. Here, influencing factors refer to potential factors or variables that may influence changes in the numerical values (i.e., parameter values) of business parameters (dependent variables). These factors are used during analysis to explain, predict, or evaluate changes in business parameters. They can be single variables or combinations of multiple variables that collectively affect business parameters. Influencing factors are typically determined based on historical experience or statistical data to identify factors or variables that may have a potential impact on changes in business parameters.
[0073] In actual implementation, for each business parameter, multiple (i.e., greater than or equal to 2) fourth influencing factors corresponding to the business parameters can be pre-set based on historical experience or data statistics. Multiple fourth influencing factors can be understood as all influencing factors that can be determined to have an impact on the business parameters.
[0074] A decision tree is a tree-like structure built based on a machine learning model, used to solve classification and regression problems. It recursively divides a dataset into smaller subsets, gradually building a tree structure consisting of a root node, intermediate nodes (also called internal nodes), and leaf nodes. Each intermediate node represents a test of a feature (here, an influencing factor); each branch represents the result of a test, and each leaf node corresponds to a feature and a predicted value (a classification result or regression value). The goal of a decision tree is to partition a dataset using a series of rules to predict or classify a target variable (here, a business parameter). The root node is the starting point of the decision tree, the leaf nodes are the end nodes, and the intermediate nodes (also called internal nodes) are the intermediate nodes in the tree, excluding the root and leaf nodes. The edges between nodes in the decision tree (including the root and intermediate nodes) correspond to decision strategies (also called splitting conditions).
[0075] In actual implementation, before constructing the first decision tree for the business parameters, a data set for fitting the first decision tree can be first constructed based on the business parameters and multiple fourth influencing factors. The data set includes the first numerical value of each fourth influencing factor at different time points within the first time period, and the third parameter value of the business parameter at the corresponding different time points.
[0076] Here, the first time period is a historical time period. In the data set, the data corresponding to the same time point (first numerical value and third parameter value) can be regarded as a sample. For example, a sample may include the third parameter value of the business parameter at time point a within the first time period and the first numerical value of each fourth influencing factor at time point a.
[0077] It is understandable that the third parameter value of the business parameter at a certain point in time is the actual value of the business parameter corresponding to that point in time. Similarly, the first value of the fourth influencing factor at a certain point in time is the actual value of the fourth influencing factor at that point in time. For example, if the business parameter is click-through rate, and the click-through rate determined at time point A is 80%, then the third parameter value of the click-through rate at time point A is 80%; similarly, if the fourth influencing factor is the advertising conversion success rate, and the advertising conversion success rate determined at time point A is 50%, then the first value of the advertising conversion success rate at time point A is 50%. In order to distinguish the parameter values corresponding to the business parameters at different time points, this application makes a distinction in the form of first, second, third, etc.; similarly, in order to distinguish the values corresponding to the fourth influencing factor at different time points, it also makes a distinction in the form of first, second, third, etc.
[0078] In actual implementation, by collecting the third parameter values and the first values corresponding to each time point in the first time period, a data set of the business parameters and the plurality of fourth influencing factors in the first time period is obtained. After obtaining the data set, each sample in the data set can be preliminarily screened to remove samples with obvious problems. For example, samples with a significant inconsistency in the potential relationship between the third parameter value and the first value can be considered problematic samples and removed.
[0079] As an example, referring to Table 1, the determined data set may be expressed as shown in Table 1:
[0080] Table 1
[0081]
[0082] In the data set shown in Table 1, Y1, Y2…Ym represent the third parameter values corresponding to the business parameter Y at different time points; X_11, X_12…X_1m represent the first values corresponding to the fourth influencing factor X1 at different time points; X_21, X_22…X_2m represent the first values corresponding to the fourth influencing factor X2 at different time points; X_n1, X_n2…X_nm represent the first values corresponding to the fourth influencing factor Xn at different time points; among them, [Y1, X_11, X_21…X_n1] can be represented as a sample in the data set.
[0083] In actual implementation, the constructed data set can be used to fit the relationship between the business parameters and multiple fourth influencing factors based on the Gradient Boosting Tree (GBT) model, so as to determine the first decision tree between the business parameters and the fourth influencing factors.
[0084] Gradient boosting is a powerful machine learning algorithm, a type of ensemble learning method. It iteratively builds multiple decision trees to gradually optimize the model's predictive performance. The core idea of gradient boosting is to minimize the loss function using gradient descent. At each step, a new decision tree is trained to fit the residual (i.e., prediction error) of the previous tree. Ultimately, the predictions from all the trees are weighted and combined to produce the final prediction.
[0085] In actual implementation, the gradient boosting tree model can be used to fit the relationship between the business parameters and multiple fourth influencing factors based on the data set, that is, the gradient boosting tree model can be trained using the data set to obtain a second decision tree that characterizes the relationship between the business parameters and multiple fourth influencing factors. Here, the second decision tree is also the decision tree that the gradient boosting tree model preliminarily fits to the business parameters and multiple fourth influencing factors. Therefore, the structure of the second decision tree may need further optimization and adjustment to obtain the final first decision tree. According to the fitting principle of the gradient boosting tree model, the number of second decision trees is usually multiple (that is, greater than or equal to 2).
[0086] In actual implementation, when fitting the second decision tree, the gradient boosting tree model typically determines a coefficient of determination for the second decision tree. The coefficient of determination is a business parameter used to evaluate the performance of the gradient boosting tree model. It measures the model's ability to explain changes in the dependent variable (such as a business parameter), that is, the extent to which the gradient boosting tree model can capture the variation in the dependent variable. The coefficient of determination ranges from [0 to 1]. Values closer to 1 indicate a better fit for the gradient boosting tree model; values closer to 0 indicate a poorer fit.
[0087] Here, the coefficient of determination of the second decision tree is used to characterize the degree of fit of the relationship between the service parameter and the plurality of fourth influencing factors.
[0088] In the above manner, the gradient boosting tree is an ensemble learning method that can gradually optimize the prediction performance of the model by iteratively constructing multiple decision trees. Therefore, each second decision tree is a component in the gradient boosting tree model, responsible for capturing the local relationship between the business parameters and the fourth influencing factor, and improving the accuracy and robustness of the overall model by combining the prediction results of multiple second decision trees.
[0089] In some embodiments, the determination coefficient of the second decision tree can be determined in the following manner, and the structure of the second decision tree can be adjusted based on the determination coefficient to obtain the first decision tree: the data set is divided into at least two sub-sample sets; for each sub-sample set, the sub-determination coefficient of the second decision tree is determined based on the sub-sample set and the second decision tree; if there is at least one sub-determination coefficient less than a first threshold, the second decision tree is iteratively fitted based on the data set until each sub-determination coefficient is greater than or equal to the first threshold, and the decision tree obtained when the fitting is stopped is used as the first decision tree.
[0090] In actual implementation, each sample in the data set can be randomly shuffled and divided into at least two sub-sample sets according to a certain ratio. For example, the data set can be divided into three sub-sample sets based on the ratio of 8:1:1. The sub-sample sets can include the main sample set, the validation set and the test set. Random shuffling and then dividing can ensure that the sample distribution in the main sample set, the validation set and the test set is basically consistent.
[0091] The primary sample set contains the largest amount of sample data and is primarily used for model training and learning patterns in the data. The validation set is used to tune hyperparameters (such as learning rate and tree depth) and select models to prevent overfitting. The test set is used to ultimately evaluate the model's generalization capabilities and ensure its performance in real-world scenarios. Therefore, by properly partitioning the data set, we can effectively train, tune, and evaluate the model, ensuring its reliability and practicality.
[0092] In actual implementation, since the data set is divided into at least two sub-sample sets, different sub-sample sets usually have different functions. Therefore, the gradient boosting tree model can be trained based on each sub-sample set to obtain the sub-determination coefficient obtained when the gradient boosting tree model is trained based on each sub-sample set.
[0093] As an example, assuming that the data set is divided into a main sample set, a test set, and a validation set, the gradient boosting tree model can be initially trained based on the main sample set to determine the first determination coefficient for fitting the relationship between the business parameter and the fourth influencing factor based on the main sample set. Then, the gradient boosting tree model is verified by the validation set, that is, the gradient boosting tree model is predicted based on the second decision tree fitted during training, and the relationship between the business parameter and the fourth influencing factor represented in the validation set is obtained, and the second determination coefficient is determined based on the first prediction result and the first label of the validation set; similarly, the gradient boosting tree model can also be tested by the test set, that is, the gradient boosting tree model is predicted based on the second decision tree fitted during training, and the relationship between the business parameter and the fourth influencing factor represented in the test set is obtained, and the third determination coefficient is determined based on the second prediction result and the second label of the test set; then, at least one of the first determination coefficient, the second determination coefficient, and the third determination coefficient is used as a sub-determination coefficient, and then, based on the three determination coefficients, it is decided how to optimize and adjust the gradient boosting tree model.
[0094] The calculation process of the coefficient of determination is described in detail below. The calculation formula of the coefficient of determination is as follows:
[0095]
[0096] Among them, SS res The Residual Sum of Squares (RSS) is the sum of squares of the errors between the model's predicted values (e.g., the first prediction, the second prediction) and the actual values (e.g., the first label, the second label). It is calculated as:
[0097]
[0098] Among them, y i is the actual value of the i-th sample; is the predicted value of the i-th sample; n is the number of samples.
[0099] In addition, SS tot The Total Sum of Squares (TSS) is the sum of the squares of the deviations between the actual values and their mean, and is calculated as:
[0100]
[0101] in, It is the mean of the third parameter values in the sample set (eg, each subsample set), for example, the mean of the third parameter values in the main sample set, the mean of the third parameter values in the test set, and the mean of the third parameter values in the validation set.
[0102] In actual implementation, a determination coefficient equal to 1 indicates that the gradient boosting tree model can perfectly fit the data, and the predicted value is completely consistent with the actual value; a determination coefficient equal to 0 indicates that the gradient boosting tree model fails to explain any changes in the business parameters; therefore, the closer the value of the determination coefficient is to 1, the better the fitting effect of the gradient boosting tree model; the closer the value is to 0, the worse the fitting effect of the gradient boosting tree model.
[0103] After calculating each sub-determination coefficient using the above formula, it is determined whether all sub-determination coefficients are greater than a first threshold. If they are all greater than the first threshold, the training of the gradient boosting tree model can be directly stopped, and the decision tree obtained at the time of stopping training is used as the first decision tree. If at least one sub-determination coefficient is less than the first threshold, the second decision tree is iteratively fitted based on the data set until all sub-determination coefficients are greater than or equal to the first threshold.
[0104] In actual implementation, if the sub-determination coefficients include a first determination coefficient, a second determination coefficient, and a third determination coefficient, the second decision tree can be iteratively fitted based on the main sample set in the data set when the first determination coefficient is greater than or equal to the first threshold and the second determination coefficient and the third determination coefficient are both less than the first threshold, and the iterative fitting is stopped until the first determination coefficient, the second determination coefficient, and the third determination coefficient are all greater than or equal to the first threshold; the decision tree obtained by fitting when the fitting is stopped is used as the first decision tree.
[0105] In actual implementation, the first threshold can be pre-set based on the actual requirements for the accuracy of the gradient boosting tree model, for example, it can be set to 0.9. If the first determination coefficient is greater than or equal to the first threshold, and the second determination coefficient and the third determination coefficient are both less than the first threshold, it means that the gradient boosting tree model still needs to be optimized. Then, the main sample set or the entire data set can be used to iteratively train the gradient boosting tree model, that is, based on the main sample set or the data set, the gradient boosting tree model is iteratively fitted to the second decision tree. After each training, the first determination coefficient, the second determination coefficient and the third determination coefficient can be re-determined in the above manner. If after a certain iterative fitting, the first determination coefficient, the second determination coefficient and the third determination coefficient are all greater than or equal to the first threshold, the iterative fitting of the second decision tree can be stopped, that is, the iterative training of the gradient boosting tree model can be stopped. Then, the latest decision tree obtained when the fitting is stopped is used as the first decision tree. The first decision tree can accurately characterize the relationship between the business parameters and multiple fourth influencing factors.
[0106] In actual implementation, if the gradient boosting tree model is trained based on the main sample set and the second decision tree is obtained, the first determination coefficient corresponding to the determined main sample set is less than the first threshold, or even too low, for example, close to 0, then it means that there is a problem with the samples in the main sample set, or there is a problem with the samples in the data set. Therefore, the estimation errors of all samples in the data set can be checked. If the estimation errors of all samples are large values, it may indicate that the quality of the fourth influencing factor in the data set is poor and cannot fully characterize the impact on changes in business parameters, or the noise in the data set is large. Therefore, the data set can be optimized by adding new samples (for example, adding new fourth influencing factors) and processing data noise; if there are some samples in the data set with large estimation errors, it may indicate that the noise in the data set is large, and the noise in the data set can be processed to exclude samples that are significantly unrelated between the first data and the third parameter value.
[0107] Here, the estimation error refers to the difference between the model's predicted value on the training sample and the actual value. It is an important indicator to measure how well the model fits the training data. When the estimation error is large, it usually means that the model fails to capture the patterns in the data well.
[0108] In actual implementation, if the gradient boosting tree model is trained based on the main sample set and the second decision tree is obtained, if the third determination coefficient corresponding to the determined test set is greater than or equal to the first threshold, the training of the gradient boosting tree model can be completed directly, that is, the second decision tree can be directly used as the first decision tree without the need for subsequent optimization and adjustment processes.
[0109] In a specific embodiment, Figure 4 This is a flow chart of the training method of the gradient boosting tree model provided in the embodiment of the present application, see Figure 4 , the training process of the gradient boosting tree model includes:
[0110] S201, determine the fourth impact factor.
[0111] A plurality of fourth influencing factors that have an impact on changes in service parameters are collected.
[0112] S202: Construct a data set (ie, a sample set).
[0113] The third parameter value of the service parameter and the first value of each fourth influencing factor at each time point in the first time period are collected to obtain a plurality of samples to form a data set.
[0114] S203, dividing the data set into a main sample set, a validation set, and a test set.
[0115] According to the preset ratio, the data set is divided into the main sample set, the verification set, and the test set.
[0116] S204: Fit the relationship between the business parameter and the fourth influencing factor using a gradient boosting tree model.
[0117] The gradient boosting tree model is trained using the main sample set to obtain a second decision tree. The first determination coefficient of the main sample set, the second determination coefficient of the validation set, and the third determination coefficient of the test set are determined. The specific determination process can be found in the above-mentioned related embodiments and will not be repeated here.
[0118] S205: Determine whether the third determination coefficient is greater than or equal to the first threshold.
[0119] If yes, the training of the gradient boosting tree model is terminated, and the fitted second decision tree is directly used as the first decision tree.
[0120] If not, proceed to S206.
[0121] S206: Determine whether the first determination coefficient is greater than or equal to a first threshold.
[0122] If yes, continue to optimize the gradient boosting tree model and jump to S204.
[0123] If not, proceed to S207.
[0124] S207, determining whether the estimation error of the sample is large.
[0125] If the estimation errors of all samples in the data set are large, the process jumps to S201 to collect the fourth influencing factor again.
[0126] If the estimation errors of some samples are large, the process jumps to S202 to perform denoising and other processing on the collected samples.
[0127] In the above manner, the fitting ability of the gradient boosting tree model can be evaluated through the first determination coefficient of the main sample set, the generalization ability of the gradient boosting tree model can be evaluated through the second determination coefficient of the validation set, and the final performance of the gradient boosting tree model can be evaluated through the third determination coefficient of the test set. In this application, these three sub-determination coefficients are used to jointly decide whether to optimize the model, which can fully understand the fitting performance of the gradient boosting tree model, avoid local optimization, ensure generalization ability, and thus improve the accuracy of the first decision tree. Moreover, when the first determination coefficient of the main sample set is small, the problems existing in the sample can also be determined by the estimated error of the sample, and improvements can be made to the sample. It can be seen that in the process of constructing the first decision tree, the performance bottlenecks and sample problems of the model can be clearly understood through the determination coefficients and estimated parameters, providing a clear direction for the optimization of the decision tree.
[0128] Continue to see Figure 3 , continue with the above step 101 for explanation.
[0129] In step 102, when the parameter value of the business parameter changes from a first parameter value to a second parameter value, based on a first decision path in the first decision tree that matches the first parameter value, a first impact value of the first influencing factor in the first decision path on the business parameter is determined, and based on a second decision path in the first decision tree that matches the second parameter value, a second impact value of the second influencing factor in the second decision path on the business parameter is determined.
[0130] In actual implementation, the parameter value of the business parameter changes from a first parameter value to a second parameter value, which usually means that the parameter value is the first parameter value at the first time point, and the parameter value becomes the second parameter value at the second time point, which means that the business parameter has changed in parameter value from the first time point to the second time point. For example, if the first time point refers to a time point on the first day of a month, the second time point can be a time point on the last day of the month, or it can be understood as the beginning of the month (the first time point) and the end of the month (the second time point).
[0131] As an example, taking the financial industry as an example, if the member net payment success amount (i.e., business parameter) for January is determined to be 1000 on the last day of January 2025, and the member net payment success amount for February is determined to be 500 on the last day of February 2025, then the member net payment success amount is used as a business parameter. If the business parameter changes from the first time point (i.e., the last day of January) to the second time point (i.e., the last day of February), then for this business parameter, the multiple fourth influencing factors that can be determined may include payment success amount, refund success amount, user login rate, advertising conversion success rate, etc., which may have a direct or potential impact on the member net payment success amount. Some factors or variables that affect the click-through rate; for example, taking the Internet industry as an example, if the click-through rate of users for a certain website is determined to be 80% on the first day of January 2025, and the click-through rate of users for a certain website is determined to be 50% on the last day of January 2025, the click-through rate can be used as a business parameter. The business parameter (i.e., click-through rate) changes from the first day of January (i.e., the first time point) to the last day of January (i.e., the second time point). Then, for the click-through rate, the multiple fourth influencing factors determined can be some factors or variables that may have a direct or potential impact on the click-through rate, such as the amount of advertising, user behavior data, website loading speed, and market competition.
[0132] In some embodiments, the first decision path can be determined in the following manner: the first decision tree takes the business parameter as the root node, and the first decision tree includes at least one node level, and each node level includes at least one influencing factor; in order from top to bottom of the node levels, based on the first parameter value and the segmentation conditions of the influencing factors in each node level, the nodes in each node level are determined in turn; the nodes in each determined node level are connected to obtain the first decision path that matches the first parameter value.
[0133] Here, the first decision tree is first described in detail.
[0134] Due to the fitting process of the gradient boosting tree model, there are usually multiple first decision trees. Each decision tree can be understood as a binary tree, including a root node, intermediate nodes (also called internal nodes), and leaf nodes. The intermediate nodes and leaf nodes constitute multiple node levels of the first decision tree. Among them, the root node is the starting point of the decision tree and contains the entire data set; the intermediate nodes and leaf nodes represent a certain fourth influencing factor; the leaf node is also the end node of the decision tree, corresponding to the predicted value for the business parameter. It can be understood that the structure of the decision tree also includes edges, which are used to connect the root node, internal nodes, and leaf nodes; the edge between two nodes can be expressed as a judgment condition or decision strategy, which can be referred to as a splitting condition here, that is, each edge corresponds to a splitting condition. Each first decision tree can divide the data set through a series of rules (based on the specific value of the fourth influencing factor) to capture the local relationship between the business parameter and the fourth influencing factor. In the gradient boosting tree, the prediction results of each first decision tree are used to gradually correct the residual (i.e., prediction error) of the previous decision tree, thereby optimizing the performance of the overall model.
[0135] As an example, Figure 5 This is a schematic diagram of the first decision tree provided in the embodiment of the present application, see Figure 5 , assuming that the gradient boosting tree model is fitted as Figure 5 The four first decision trees (a), (b), (c), and (d) shown in Figure 5 The first decision tree shown in (a) is used as an example to illustrate. The first decision tree includes a root node and two node levels. Among them, in the order of the node levels from top to bottom, the first node level includes two internal nodes, node 1-1 and node 1-2, and the second node level includes four leaf nodes, node 1-11, node 1-12, node 1-21, and node 1-22.
[0136] As an example, in the first decision tree, two nodes belonging to the same parent node usually correspond to the same fourth impact factor, for example, see Figure 5In the first decision tree shown in Figure (a), nodes 1-1 and 1-2 can correspond to the same fourth influencing factor. The splitting condition corresponding to the edge connecting node 1-1 and the root node, and the splitting condition corresponding to the edge connecting node 1-2 and the root node, are both set for the fourth influencing factor corresponding to node 1-1 and node 1-2. For example, if the fourth influencing factor corresponding to nodes 1-1 and 1-2 is the advertising conversion success rate, the splitting condition for the edge between node 1-1 and the root node can be that the advertising conversion success rate is greater than 70%, while the splitting condition for the edge between node 1-2 and the root node can be that the advertising conversion success rate is less than 70%. Similarly, if the fourth influencing factor corresponding to leaf nodes 1-11 and 1-12 is the advertising delivery volume, the splitting condition for the edge between leaf nodes 1-11 and 1-1 can be that the advertising delivery volume is greater than 1000, while the splitting condition for the edge between leaf nodes 1-12 and 1-1 can be that the advertising delivery volume is less than 1000.
[0137] It is understandable that Figure 5 Only a schematic diagram of the first decision tree is given. The tree depths of different first decision trees may be different. Moreover, different first decision trees may involve the same fourth influencing factor or different fourth influencing factors. For example, Figure 5 The fourth influencing factor involved in the first decision tree shown in Figure (a) can be Figure 5 The fourth influencing factors involved in the first decision tree shown in Figure (b) are partially the same, completely different, or completely the same.
[0138] In actual implementation, after the first decision tree of the business parameters and the fourth influencing factors is obtained by fitting the gradient boosting tree model, the relationship between the business parameters and the fourth influencing factors can be expressed by Y = F(X_1, X_2, ...., X_n), where Y represents the business parameters, X_1, X_2, ...., X_n represent n fourth influencing factors, and F() represents the gradient boosting tree model.
[0139] Furthermore, based on fitting the first decision tree with the gradient boosting tree model, the expression between the business parameter and the fourth influencing factor can also be:
[0140]
[0141] Where k represents the number of the first decision trees; q k (x) represents the number of the leaf node to which sample x belongs in the kth first decision tree; It represents the predicted value (i.e., influence value) corresponding to the leaf node to which the sample x belongs in the kth first decision tree.
[0142] During actual implementation, since the business parameters change from the first time point to the second time point, it is necessary to perform analysis based on the actual data of the business parameters and the fourth influencing factors at the first time point and the actual data at the second time point to determine which changes in the fourth influencing factors have affected the business parameters and caused the business parameters to change.
[0143] Therefore, when determining the first decision path matching the first parameter value, it is also necessary to obtain the second value of each fourth influencing factor at the first time point, and then use the first parameter value and the second values of the plurality of fourth influencing factors as the sample x b The input is input into the trained gradient boosting tree model, so that the gradient boosting tree model determines a first decision path matching the first parameter value in the first decision tree, and predicts the first impact value based on the first decision path.
[0144] In actual implementation, when the gradient boosting tree model determines the first decision path that matches the first parameter value through the first decision tree, it actually substitutes the second value corresponding to the fourth influencing factor at the first time point into the first decision tree, determines the nodes involved in the second value in the first decision tree, and then connects the involved nodes to obtain the first decision path that matches the first parameter value and the second value.
[0145] As an example, Figure 5 (a) is used as an example for explanation. It is assumed that the fourth influencing factor corresponding to the internal nodes 1-1 and the internal nodes 1-2 in the first decision tree shown in Figure (a) is the advertising conversion success rate. The segmentation condition corresponding to the internal node 1-1 can be that the advertising conversion success rate is greater than 70%, and the segmentation condition corresponding to the internal node 1-2 can be that the advertising conversion success rate is less than 70%; the fourth influencing factor corresponding to the leaf nodes 1-11 and the leaf nodes 1-12 is the advertising delivery volume. The segmentation condition corresponding to the leaf node 1-11 can be that the advertising delivery volume is greater than 1000, and the segmentation condition corresponding to the leaf node 1-12 can be that the advertising delivery volume is less than 1000; the fourth influencing factor corresponding to the leaf nodes 1-21 and the leaf nodes 1-22 is the user behavior data. The segmentation condition corresponding to the leaf node 1-21 can be that the user behavior data meets the behavioral expectations, and the segmentation condition corresponding to the leaf node 1-22 can be that the user behavior data does not meet the behavioral expectations.
[0146] Assume that the second values of the plurality of fourth influencing factors at the first time point include: the second value corresponding to the advertising conversion success rate is 75%, and the second value corresponding to the advertising delivery volume is 2000. Figure 5The first decision tree shown in Figure (a) determines the nodes to which the second value is adapted in each node level in sequence according to the influencing factors involved in the nodes in each node level and the segmentation conditions corresponding to the nodes, in order from top to bottom. For example, the first parameter value of the business parameter corresponds to the root node of the first decision tree. According to the order from top to bottom, when passing through the first node level (including node 1-1 and node 1-2), according to "the second value corresponding to the advertising conversion success rate is 75%", it can be determined that the second value matches the fourth influencing factor and segmentation condition corresponding to node 1-1. Therefore, according to the second value in the first The node determined in the node hierarchy is the internal node 1-1; continue to judge downward, because the node that matches the first parameter value in the first node hierarchy is node 1-1, therefore, when judging the second node hierarchy, only the child nodes of node 1-1 in the next node hierarchy (i.e., nodes 1-11 and 1-12) need to be considered. According to "the second numerical value corresponding to the advertising delivery volume is 2000", it can be determined that the second numerical value matches the fourth influencing factor and the splitting condition corresponding to node 1-11 in the second node hierarchy. Therefore, the node determined in the second node hierarchy according to the second numerical value is node 1-11. Based on the above process, the nodes determined in each node hierarchy in the first decision tree using the second numerical value are node 1-1 and node 1-11. Therefore, by connecting the root node, node 1-1, and node 1-11, the first decision path that matches the first parameter value can be obtained.
[0147] It can be understood that the fourth impact factor corresponding to each node in the first decision path is the first impact factor in the first decision path.
[0148] In some embodiments, after determining the first decision path, Figure 3 The "determining the first impact value of the first influencing factor in the first decision path on the business parameter based on the first decision path matching the first parameter value in the first decision tree" in step 102 shown can be achieved in the following way: for each node level, based on the nodes in the node level located in the first decision path and the first impact factors of the nodes in the first decision path, determine the second impact value of the first impact factor in the node level; add the second impact values of the first impact factor in each node level to obtain the first impact value of the first impact factor in the first decision path on the business parameter.
[0149] In actual implementation, the mapping relationship between the node position and the impact value on the decision tree can be pre-configured for the decision tree. Figure 5(a) is used as an example. If the pre-configured node position is the root node with an influence value of 10, the node position is 1-1 with an influence value of 20, the node position is 1-2 with an influence value of 10, the node position is 1-11 with an influence value of 30, the node position is 1-12 with an influence value of 20, the node position is 1-21 with an influence value of 30, and the node position is 1-22 with an influence value of 10. Continuing with the previous example, after determining the first decision path, according to the above mapping relationship, it can be determined that the influence value of the root node involved in the first decision path is 10, the influence value of node 1-1 is 20, and the influence value of node 1-11 is 30. Furthermore, it can be determined that the second influence value corresponding to the first influence factor (advertising conversion success rate) corresponding to node 1-1 is 20, and the second influence value corresponding to the first influence factor (advertising delivery volume) corresponding to node 1-11 is 30. Then, the second influence value 20 of the first influence factor (advertising conversion success rate) and the second influence value 30 of the first influence factor (advertising delivery amount) in the first decision path are added together to obtain the first influence value 50 of the first influence factor on the business parameter in the first decision path. In some cases, the influence value of the root node in the first decision path and the second influence value of the first influence factor in each node level can also be added together to obtain the first influence value of the first influence factor on the business parameter in the first decision path.
[0150] Here, the first influence value calculated based on the above method can be understood as the gradient boosting tree model based on the first parameter value and the second value at the first time point. Figure 5 (a) The predicted values of the first decision tree in Figure .
[0151] As mentioned above, the number of first decision trees fitted by the gradient boosting tree model is usually multiple. Figure 5 As an example of the first decision tree shown in FIG (a), a specific process of determining the first decision path in the first decision tree and the first impact value of the first influencing factor on the business parameter in the first decision path is given. Other first decision trees (such as Figure 5 The first decision tree shown in (b), (c), and (d) Figure 5 The determination method of the first decision tree in Figure (a) is consistent.
[0152] In actual implementation, the sum of the first influence values of each first decision tree can be used as the predicted value of the first decision tree by the gradient boosting tree model based on the first parameter value and the second value at the first time point.
[0153] As an example, based on the expression between the above-mentioned business parameter and the fourth influencing factor, after substituting the first parameter value and the second value at the first time point into the first decision tree, the expression between the business parameter at the first time point and the fourth influencing factor can be expressed as follows:
[0154]
[0155] Among them, b refers to the first time point, y b Indicates the business parameters at the first time point.
[0156] In actual implementation, if there is one first decision tree, the predicted value corresponding to the first decision tree is the first impact value corresponding to the first decision tree, that is, the first impact value of the gradient boosting tree model on the business parameter prediction at the first time point for the first decision tree. If there are multiple first decision trees, the first impact value of the gradient boosting tree model on the business parameter prediction at the first time point is the sum of the predicted values of all first decision trees, that is, the first impact value of the gradient boosting tree model on the business parameter prediction at the first time point is the sum of the first impact values (i.e., predicted values) corresponding to all first decision trees.
[0157] Similarly, in some embodiments, the second decision path can be determined in the following manner: according to the order of the node hierarchy from top to bottom, based on the second parameter value and the segmentation conditions of the influencing factors in each node hierarchy, the nodes in each node hierarchy are determined in turn; the nodes in each determined node hierarchy are connected to obtain a second decision path that matches the second parameter value.
[0158] In actual implementation, the second decision path is determined in the same manner as the first decision path, except that the second parameter value at the second time point and the third values of the plurality of fourth influencing factors are substituted. That is, the second decision path is determined based on the second parameter value and the third values. Therefore, the process of determining the second decision path is not further described here.
[0159] As an example, following the above example Figure 5To illustrate the hypothetical influencing factors and splitting conditions in (a), if the third values of the multiple fourth influencing factors at the second time point include: the third value of the advertising conversion success rate is 55%, and the user behavior data does not meet behavioral expectations. Then, according to the order of the node levels from top to bottom, it can be determined that the third value matches the fourth influencing factor and splitting condition corresponding to node 1-2 based on "the third value of the advertising conversion success rate is 55%". Therefore, the node determined in the first node level based on the third value is internal node 1-2; continuing to judge downward, since the node determined in the first node level is internal node 1-2, when judging the second node level, only the child nodes of node 1-2 in the next node level (i.e., node 1-21 and node 1-22) need to be considered. Therefore, according to "the user behavior data does not meet behavioral expectations", it can be determined that the third value matches the fourth influencing factor and splitting condition corresponding to node 1-22 in the second node level. Therefore, the node determined in the second node level based on the third value is node 1-22. Based on the above process, the nodes determined in each node level in the first decision tree using the third value are node 1-2 and node 1-22. Therefore, by connecting the root node, node 1-2 and node 1-22, a second decision path matching the second parameter value can be obtained.
[0160] It can be understood that the fourth influencing factor corresponding to each node in the second decision path is the second influencing factor in the second decision path.
[0161] In some embodiments, after determining the second decision path, Figure 3 The "determining the second impact value of the second influencing factor in the second decision path on the business parameter based on the second decision path matching the second parameter value in the first decision tree" in step 102 shown can be achieved in the following way: for each node level, based on the nodes in the node level located in the second decision path and the second impact factors of the nodes in the second decision path, determine the fourth impact value of the second impact factor in the node level; add the fourth impact value of the second impact factor in each node level to obtain the first impact value of the second impact factor in the second decision path on the business parameter.
[0162] As an example, if, based on a mapping relationship between node positions and influence values on a decision tree pre-configured for the decision tree, the influence value corresponding to the root node is determined to be 10, the influence value corresponding to node 1-2 is 20, and the influence value corresponding to node 1-22 is 40, then the fourth influence value of the second influence factor (advertising conversion success rate) corresponding to node 1-2 can be determined to be 20, and the fourth influence value of the second influence factor (user behavior data) corresponding to node 1-22 can be determined to be 40. Therefore, the fourth influence value 20 of the second influence factor (advertising conversion success rate) and the fourth influence value 40 of the second influence factor (user behavior data) can be added together to obtain the second influence value of the second influence factor on the business parameter in the second decision path.
[0163] Similarly, if there are multiple first decision trees, after substituting the second parameter value and the third value at the second time point into the first decision tree, the expression between the business parameter at the second time point and the fourth influencing factor can be expressed as:
[0164]
[0165] Among them, e refers to the second time point, y e Indicates the service parameters at the second time point.
[0166] It can be seen that if there are multiple first decision trees, the second impact value of the gradient boosting tree model on the business parameter prediction at the second time point is the sum of the prediction values of all first decision trees, that is, the second impact value of the gradient boosting tree model on the business parameter prediction at the second time point is the sum of the second impact values (i.e., prediction values) corresponding to all first decision trees.
[0167] Continue to see Figure 3 , continue with the above step 102 for description.
[0168] In step 103, based on the first impact value and the second impact value, a third impact factor for causing a change in the parameter value of the service parameter is determined from the first impact factor and the second impact factor.
[0169] In some embodiments, Figure 3 Step 103 shown can be implemented in the following manner: when the first impact value is different from the second impact value, based on the difference between the first decision path and the second decision path, determine the fifth impact factor that affects the change in the parameter value of the business parameter among the first impact factor and the second impact factor; based on the first impact value and the second impact value, determine the third impact value of each fifth impact factor; and determine the fifth impact factor whose third impact value is greater than the second threshold as the third impact factor of the change in the parameter value of the business parameter.
[0170] In actual implementation, if the first impact value predicted by the first decision tree is different from the second impact value, it means that the first decision path and the second decision path must be different, that is, there must be different nodes in the first decision path and the second decision path, and the impact factors corresponding to these different nodes can be regarded as the fifth impact factor that affects the change in the parameter value of the business parameter.
[0171] As an example, based on the expressions at the first time point and the second time point given in the above embodiment, the difference between the service parameters at the first time point and the second time point can be schematically expressed by the following formula:
[0172]
[0173] It should be noted that here It does not directly refer to the difference in the prediction values of the first decision tree, but rather represents the difference in the decision path of the first decision tree at the first time point and the second time point.
[0174] In some embodiments, the above-mentioned "based on the difference between the first decision path and the second decision path, determining the fifth influencing factor that affects the change in the parameter value of the business parameter in the first influencing factor and the second influencing factor" can be achieved in the following way: determining multiple different nodes in the first decision path and the second decision path, and the parent nodes of the multiple nodes; using the influencing factors corresponding to the multiple nodes and the influencing factors corresponding to the parent nodes as the fifth influencing factor.
[0175] In actual implementation, since there can be multiple first decision trees and the fourth influencing factors involved in each first decision tree may be different, the difference comparison can be split into each first decision tree. Based on the difference between the first decision path at the first time point and the second decision path at the second time point on each first decision tree, the fifth influencing factor that affects the business parameters in each first decision tree and the third impact value of the fifth influencing factor on the business parameters in the first decision tree are determined. Finally, the fifth influencing factors that affect the business parameters determined on all first decision trees are integrated to obtain the third impact value of each fifth influencing factor on the business parameters.
[0176] Below, for each first decision tree, a method for determining the fifth influencing factor in each first decision tree is explained.
[0177] As an example, Figure 6A This is a first schematic diagram of a method for determining the fifth impact factor provided in an embodiment of the present application, see Figure 6A , Figure 6AThe tree structure of the first decision tree is schematically shown, and the first decision tree includes a root node, 6 internal nodes (node 1-1, node 1-2, node 1-11, node 1-12, node 1-21, node 1-22), and 8 leaf nodes (node 1, node 2, node 3, node 4, node 5, node 6, node 7, and node 8). Assuming that the first decision path determined based on the first parameter value and the second numerical value is obtained by connecting the root node, node 1-2, node 1-21, and node 5, and the second decision path determined based on the second parameter value and the third numerical value is obtained by connecting the root node, node 1-2, node 1-21, and node 6, it can be determined that the different multiple nodes in the first decision path and the second decision path include node 5 and node 6, and the parent node of node 5 and node 6 is node 1-21. Therefore, the impact factor corresponding to node 5, the impact factor corresponding to node 6, and the impact factor corresponding to node 1-21 can be used as the fifth impact factor.
[0178] It can be understood that the fifth influencing factor belongs to the first influencing factor in the first decision path and the second influencing factor in the second decision path, while the first influencing factor and the second influencing factor both belong to the fourth influencing factor. Figure 6A , node 5 is a node in the first decision path, so node 5 corresponds to the first influencing factor; node 6 is a node in the second decision path, so node 6 corresponds to the second influencing factor; and nodes 1-21 belong to both the first decision path and the second decision path, so the influencing factors corresponding to nodes 1-21 can be regarded as both the first influencing factor and the second influencing factor.
[0179] As an example, Figure 6B This is a second schematic diagram of the method for determining the fifth impact factor provided in the embodiment of the present application, see Figure 6B , Figure 6B The tree structure of the first decision tree is shown with Figure 6A The tree structures shown are consistent and can all refer to the tree structure of the first decision tree. Assuming that the first decision path determined based on the first parameter value and the second numerical value is obtained by connecting the root node, node 1-2, node 1-21, and node 6, and the second decision path determined based on the second parameter value and the third numerical value is obtained by connecting the root node, node 1-2, node 1-22, and node 7, it can be determined that the different multiple nodes in the first decision path and the second decision path include node 6, node 7, node 1-21, and node 1-22, and the parent node of the different multiple nodes is node 1-2. Therefore, the impact factor corresponding to node 6, the impact factor corresponding to node 7, the impact factor corresponding to node 1-21, and the impact factor corresponding to node 1-22 can be used as the fifth impact factor.
[0180] In some embodiments, the above-mentioned "determining the third influence value of each fifth influence factor based on the first influence value and the second influence value" can be achieved in the following way: if multiple different nodes belong to the same node level, determine the difference between the second influence value and the first influence value, and the average of the second influence value and the first influence value; use the difference as the third influence value of the fifth influence factor corresponding to the parent node, and use the average value as the third influence value of the fifth influence factor corresponding to each node.
[0181] After the fifth impact factor on each first decision tree is determined in the above manner, the third impact value of each fifth impact factor may be calculated based on the first impact value and the second impact value.
[0182] In actual implementation, there are certain differences in the calculation method of the third impact value of the fifth impact factor when different nodes of the first decision path and the second decision path belong to the same node level or belong to different node levels.
[0183] In actual implementation, if multiple nodes belong to the same node level, continue to refer to Figure 6A , the different nodes of the first decision path and the second decision path are node 5 and node 6, and node 5 and node 6 belong to the same node level. In this case, the difference between the second influence value and the first influence value can be used as the third influence value of the fifth influence factor corresponding to the parent node of node 5 and node 6, and the average of the second influence value and the first influence value can be used as the third influence value of the fifth influence factor corresponding to node 5 and the third influence value of the fifth influence factor corresponding to node 6.
[0184] If the difference calculated here and the subsequent difference calculations are negative, the absolute value needs to be taken.
[0185] Understandably, continue to see Figure 6A Since node 5 and node 6 belong to the same parent node, the fifth influence factor corresponding to node 5 and the fifth influence factor corresponding to node 6 are actually the same influence factor. Assuming that node 5 and node 6 both correspond to the fifth influence factor of user behavior data, after determining the third influence value of user behavior data for node 5 and the third influence value of user behavior data for node 6, the third influence value calculated for node 5 and the third influence value calculated for node 6 can be added together as the third influence value of user behavior data in the first decision tree.
[0186] In some embodiments, different multiple nodes include a first leaf node, a second leaf node and an intermediate node; if different multiple nodes belong to different node levels, the above-mentioned "determining the third influence value of each fifth influence factor based on the first influence value and the second influence value" can also be achieved in the following way: for the first leaf node, determine the first influence value of the first leaf node and the average of the influence values of the adjacent leaf nodes of the first leaf node, and use the average as the third influence value of the fifth influence factor corresponding to the first leaf node; for the second leaf node, determine the second influence value of the second leaf node and the average of the influence values of the adjacent leaf nodes of the second leaf node, and use the average as the third influence value of the fifth influence factor corresponding to the second leaf node; for the intermediate node, respectively determine the influence values of the influence factors corresponding to the two child nodes of the intermediate node, and determine the difference in the influence values of the influence factors corresponding to the two child nodes, and determine the third influence value of the fifth influence factor corresponding to the intermediate node based on the difference.
[0187] In actual implementation, if different nodes of the first decision tree and the second decision tree belong to different node levels, the calculation methods of the third impact values of the fifth impact factors corresponding to the leaf nodes, intermediate nodes and parent nodes are different.
[0188] In actual implementation, when there are differences between the first decision path and the second decision path, the leaf nodes that the two paths must fall into are different. Therefore, the different nodes between the first decision path and the second decision path must include the first leaf node in the first decision path and the second leaf node in the second decision path, as well as the intermediate node to which the first leaf node belongs and the intermediate node to which the second leaf node belongs. It can be understood that if the intermediate node to which the first leaf node belongs and the intermediate node to which the second leaf node belongs are the same node, then the intermediate node is the parent node of the two different leaf nodes, which is equivalent to the situation in the above embodiment where different nodes belong to the same node level; if the intermediate node to which the first leaf node belongs and the intermediate node to which the second leaf node belongs are not the same node, then the different nodes between the first decision path and the second decision path include at least the first leaf node, the intermediate node to which the first leaf node belongs, the second leaf node, the intermediate node to which the second leaf node belongs, and the parent node to which the two intermediate nodes belong, that is, the situation where different nodes on the two paths belong to different node levels.
[0189] As an example, see Figure 6BThe first decision path determined based on the first parameter value and the second numerical value is obtained by connecting the root node, node 1-2, node 1-21, and node 6. The second decision path determined based on the second parameter value and the third numerical value is obtained by connecting the root node, node 1-2, node 1-22, and node 7. It can be seen that the different nodes between the first decision path and the second decision path include node 6 (first leaf node), node 7 (second leaf node), node 1-21 (intermediate node), and node 1-22 (intermediate node). It can be seen that nodes 6 and 7 belong to the same node level, and nodes 1-21 and 1-22 belong to the same node level, but these two node levels are not the same level.
[0190] Based on this, for the fifth influence factor corresponding to node 6 (the first leaf node), the influence value corresponding to the adjacent leaf node (node 5) of node 6 can be determined based on the mapping relationship between the node position and the influence value configured above. Then, the first influence value corresponding to node 6 (the first leaf node) and the influence value corresponding to node 5 (the adjacent leaf node) are summed and averaged to obtain the average value, and the obtained average value is used as the third influence value of the fifth influence factor corresponding to node 6. Here, the adjacent leaf node (node 5) of node 6 (the first leaf node) is a leaf node that belongs to the same intermediate node as node 6. Similarly, for the fifth influence factor corresponding to node 7 (the second leaf node), the influence value corresponding to the adjacent leaf node (node 8) of node 7 can be determined based on the mapping relationship between the node position and influence value configured above. Then, the second influence value corresponding to node 7 (the second leaf node) and the influence value corresponding to node 8 (the adjacent leaf node) are summed and averaged to obtain the average value, and the obtained average value is used as the third influence value of the fifth influence factor corresponding to node 7. Here, the adjacent leaf node (node 8) of node 7 (the second leaf node) is a leaf node that belongs to the same intermediate node as node 7.
[0191] Continuing from the above Figure 6BThe following example illustrates how to determine the third influence value of the fifth influence factor corresponding to the intermediate node. Taking the intermediate node (node 1-21) as an example, the two child nodes of the intermediate node (node 1-21) are nodes 5 and 6. As shown above, the influence value corresponding to node 5 (denoted as w5) and the first influence value corresponding to node 6 (denoted as w6) are already determined. Therefore, the difference (w5-w6) between the influence value corresponding to node 5 and the first influence value corresponding to node 6 can be determined, and 1 / 2 of this difference can be used as the third influence value of the fifth influence factor corresponding to the intermediate node (node 1-21) (i.e., (w5-w6) / 2). In the same way, the third influence value of the fifth influence factor corresponding to the intermediate node (node 1-22) can be calculated as (w7-w8) / 2, where w7 represents the second influence value corresponding to node 7 and w8 represents the influence value corresponding to node 8.
[0192] Here, (w5-w6) / 2 represents the third influence value of the fifth influence factor corresponding to the intermediate node (node 1-21). The influence value corresponding to the intermediate node (node 1-21) can be the average of the influence values of nodes 5 and 6, that is, (w5+w6) / 2. Similarly, the influence value corresponding to the intermediate node (node 1-22) can be the average of the influence values of nodes 7 and 8, that is, (w7+w8) / 2. Similar to the calculation method of the third influence value of the fifth influence factor corresponding to the intermediate node, the third influence value of the fifth influence factor of the parent node (node 1-2) can be determined based on the difference in influence values corresponding to the two child nodes of the parent node. For example, the two child nodes of the parent node (node 1-2) are determined to be nodes 1-21 and 1-22. Then, the difference in influence values of the two child nodes is determined to be (w7+w8) / 2-(w5+w6) / 2, and 1 / 2 of this difference is used as the third influence value of the fifth influence factor corresponding to the parent node (node 1-2).
[0193] It is understandable that if different intermediate nodes of the first decision tree and the second decision tree belong to different node levels, the calculation method of the third influence value of the fifth influence factor corresponding to each intermediate node and parent node is the same as above.
[0194] In some cases, in addition to the situation where the positions of the first leaf node and the second leaf node are different, the positions of the first leaf node and the second leaf node may be the same. In this case, the first decision path and the second decision path are usually consistent, and the first impact value and the second impact value are usually the same. Therefore, in this case, there is no fifth impact factor in the first decision tree that affects the change of business parameters.
[0195] Based on the above situations, the fifth influencing factors involved in each first decision tree can be determined, and the third influence values corresponding to each fifth influencing factor on each first decision tree can be determined.
[0196] In actual implementation, if there is an identical fifth influencing factor among the fifth influencing factors of multiple first decision trees, the corresponding third influencing values of these identical fifth influencing factors on each first decision tree can be added together to obtain a third influencing value used to characterize the degree of influence of the fifth influencing factor on the business parameters.
[0197] As an example, if four first decision trees are fitted, and it is determined that there is one fifth influencing factor in the first first decision tree (assuming it is the advertising conversion success rate), two fifth influencing factors in the second first decision tree (assuming they include the advertising conversion success rate and user behavior data), one fifth influencing factor in the third first decision tree (assuming it is user behavior data), and three fifth influencing factors in the fourth first decision tree (assuming they include user behavior data, website loading speed, and advertising delivery volume), then for the fifth influencing factor of the advertising conversion success rate, the third influencing value of the advertising conversion success rate calculated in the first first decision tree and the fifth influencing value of the advertising conversion success rate calculated in the second first decision tree can be used as the influencing factors. The sum of the third influence values of the rate is used as the third influence value of the fifth influencing factor, the advertising conversion success rate. The third influence value obtained after the sum represents the degree of influence of the fifth influencing factor, the advertising conversion success rate, on the change of business parameters. The larger the third influence value, the greater the influence degree. Similarly, for the fifth influencing factor, the user behavior data, the sum of the third influence value of the user behavior data calculated in the second first decision tree, the third influence value of the user behavior data calculated in the third first decision tree, and the third influence value of the user behavior data calculated in the fourth first decision tree can be used as the third influence value of the fifth influencing factor, the user behavior data. The same applies to the others and will not be repeated here.
[0198] Through the above method, the third influence value of each fifth influence factor is predicted based on the gradient boosting tree model, that is, the third influence value of each fifth influence factor is calculated based on machine learning, which no longer relies on the precise mathematical expression proposed in the relevant technology, can improve accuracy and reliability, so that the third influence factor can be accurately captured, so that the business data processing method of this application can be widely used in different industries and complex problems, and the versatility of the data analysis method is improved.
[0199] Based on the above method, the fifth influencing factor that affects the change of the parameter value of the business parameter can be determined from multiple first decision trees, and the third influence value of each fifth influencing factor can be calculated. Since the larger the third influence value, the greater the impact of the fifth influencing factor on the change of the business parameter, the fifth influencing factor whose third influence value is greater than the second threshold can be determined as the key influencing factor that causes the change of the business parameter, that is, as the third influencing factor.
[0200] In actual implementation, after determining the third influencing factor that has the main impact on the change in the parameter value of the business parameter, we can start from the third influencing factor, analyze the reasons for the change in the business parameter, and make corresponding plans and adjustments to the subsequent third influencing factors, so that the business parameters will meet their indicator requirements in the future.
[0201] As an example, in the risk control system, after determining the third influencing factor, a risk resolution strategy that is compatible with the third influencing factor can be determined, and the subsequent operation of the first business can be planned and adjusted based on the risk resolution strategy, so that the business parameters can meet the indicator requirements, thereby stabilizing the operating status of the first business, thereby improving the stability and security of the risk control system when monitoring the first business.
[0202] In the above manner, when performing data analysis on the first business, a business parameter reflecting the operating status of the first business is determined, and a first decision tree corresponding to the business parameter is determined. When the parameter value of the business parameter changes from a first parameter value to a second parameter value, a first decision path matching the first parameter value and a second decision path matching the second parameter value can be determined in the first decision tree, thereby determining a first impact value of the first influencing factor in the first decision path on the business parameter and a second impact value of the second influencing factor in the second decision path on the business parameter. Finally, based on the first impact value and the second impact value, the parameter value of the business parameter can be determined from the first impact factor and the second impact factor. The third influencing factor that changes can be used in this way to automatically capture the nonlinear relationship and interaction between business parameters and influencing factors using a decision tree. In this way, when analyzing changes in business parameters of the first business, there is no need to preset precise mathematical expressions between business parameters and influencing factors as in the related art. When business parameters change, the third influencing factor that causes the business parameters to change can be accurately captured based on the first decision tree, thereby being able to analyze changes in business parameters based on the third influencing factor. This method of processing business data can be applicable to different business scenarios and different data characteristics, thereby improving the flexibility and versatility of data processing methods in practical applications.
[0203] In a specific embodiment, taking the risk control system performing risk monitoring on consumer finance business as an example, when the risk control system detects a change in the net successful payment amount of a member (ie, a business parameter), the business parameter needs to be analyzed.
[0204] Figure 7 This is a schematic diagram of the corresponding service parameters and the fourth impact factor provided in the embodiment of the present application, see Figure 7 In the related art, it can support attribution analysis with fixed mathematical expressions. For example, in the first-level formula decomposition, the addition and subtraction decomposition method can be used to calculate the degree of influence of the fourth influencing factor on the business parameters, such as the net successful payment amount of members = successful payment amount - successful refund amount; in the second-level formula decomposition, the addition and subtraction or multiplication and division decomposition method can be used to calculate the degree of influence of the fourth influencing factor on the business parameters, such as successful payment amount = successful activation payment amount + successful renewal payment amount, successful refund amount = number of successful refund members * average number of refund orders per person * refund amount per order.
[0205] The business data processing method provided by the embodiment of the present application can support the analysis of the impact of the fourth influencing factor without a definite relationship (i.e., without an exact expression) on the business parameters. In this way, the impact of the upstream and downstream related influencing factors on the business parameters is often very important, which can achieve deeper root cause mining. Here, the fourth influencing factor without a definite relationship can be determined based on business experience analysis, such as Figure 7 The fourth influencing factors shown in the figure include user login rate, advertising conversion success rate, user withdrawal success rate, new user conversion rate, old user churn rate, equity penetration rate, equity collection rate, and equity utilization rate.
[0206] When analyzing changes in business parameters using the business data processing method of an embodiment of the present application, the relevant fourth influencing factor with a precise data expression determined in the relevant technology and the fourth influencing factor with no definite relationship determined based on empirical analysis can both be used as the fourth influencing factor for establishing the first decision tree with the business parameters.
[0207] As an example, by using the method in the above embodiment, at least one first decision tree can be obtained by fitting the business parameters and the fourth influencing factor. Figure 8 This is a schematic diagram of the first decision tree of the service parameters and the fourth influencing factor provided in the embodiment of the present application, see Figure 8 , Figure 8 The schematic diagram shows one of the first decision trees of the business parameters and the fourth influencing factors, wherein the first decision tree shown involves the above-mentioned six fourth influencing factors. In the case where there are multiple first decision trees, the tree depths of the other first decision trees can be the same as Figure 8The tree depths of the first decision trees shown are different, and the fourth influencing factors involved in different first decision trees may be partially different.
[0208] Subsequently, the fifth influencing factor that affects the change of business parameters can be determined from the fourth influencing factors involved in each first decision tree, and after calculating the third influencing value of each fifth influencing factor, the third influencing factor that has a major impact on the change of business parameters can be screened out from the fifth influencing factors based on the second threshold, and the corresponding risk resolution strategy can be determined based on the third influencing factor, so as to make adjustments and plans for the subsequent operation of the consumer finance business based on the risk resolution strategy, that is, to take certain optimization measures for the third influencing factor based on the risk resolution strategy, so as to keep the operating status of the consumer finance business in a better direction.
[0209] The following continues to describe the exemplary structure of the business data processing device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the service data processing device 455 of the memory 450 may include:
[0210] A first determining module 4551 is configured to determine a service parameter reflecting an operating state of a first service and a first decision tree corresponding to the service parameter, wherein the first decision tree includes multiple decision paths;
[0211] The second determination module 4552 is used to determine, when the parameter value of the business parameter changes from the first parameter value to the second parameter value, based on the first decision path in the first decision tree that matches the first parameter value, the first impact value of the first influencing factor in the first decision path on the business parameter, and to determine, based on the second decision path in the first decision tree that matches the second parameter value, the second impact value of the second influencing factor in the second decision path on the business parameter.
[0212] The third determining module 4553 is configured to determine, based on the first impact value and the second impact value, a third impact factor causing a change in the parameter value of the service parameter from among the first impact factor and the second impact factor.
[0213] In some embodiments, the first determination module 4551 is also used to construct a data set based on a plurality of fourth influencing factors preset for constructing a first decision tree, the data set including the first numerical value of each fourth influencing factor at different time points, and the third parameter value of the business parameter at different time points; based on the data set, the relationship between the business parameters and the plurality of fourth influencing factors is fitted to obtain a second decision tree; the determination coefficient of the second decision tree is determined, and the structure of the second decision tree is adjusted based on the determination coefficient to obtain a first decision tree.
[0214] In some embodiments, the first determination module 4551 is also used to divide the data set into at least two sub-sample sets; for each sub-sample set, based on the sub-sample set and the second decision tree, determine the sub-determination coefficient of the second decision tree; if there is at least one sub-determination coefficient less than the first threshold, iteratively fit the second decision tree based on the data set until each sub-determination coefficient is greater than or equal to the first threshold, and the decision tree obtained when the fitting is stopped is used as the first decision tree.
[0215] In some embodiments, the first decision tree takes the business parameter as the root node, and the first decision tree includes at least one node level, and each node level includes at least one influencing factor; the second determination module 4552 is also used to determine the nodes in each node level in sequence according to the order of the node levels from top to bottom, based on the first parameter value and the segmentation conditions of the influencing factors in each node level; connect the nodes in each determined node level to obtain a first decision path that matches the first parameter value.
[0216] In some embodiments, the second determination module 4552 is also used to determine, for each node level, the second impact value of the first impact factor in the node level based on the nodes in the node level located in the first decision path and the first impact factors of the nodes in the first decision path; and add the second impact values of the first impact factor in each node level to obtain the first impact value of the first impact factor in the first decision path on the business parameter.
[0217] In some embodiments, the third determination module 4553 is also used to determine, when the first impact value is different from the second impact value, based on the difference between the first decision path and the second decision path, a fifth impact factor that affects the change in the parameter value of the business parameter among the first impact factor and the second impact factor; determine the third impact value of each fifth impact factor based on the first impact value and the second impact value; and determine the fifth impact factor whose third impact value is greater than the second threshold as the third impact factor of the change in the parameter value of the business parameter.
[0218] In some embodiments, the third determination module 4553 is further used to determine multiple different nodes in the first decision path and the second decision path, as well as the parent nodes of the multiple nodes; and use the impact factors corresponding to the multiple nodes and the impact factors corresponding to the parent nodes as the fifth impact factor.
[0219] In some embodiments, the third determination module 4553 is also used to determine the difference between the second influence value and the first influence value, and the average of the second influence value and the first influence value if multiple different nodes belong to the same node level; use the difference as the third influence value of the fifth influence factor corresponding to the parent node, and use the average as the third influence value of the fifth influence factor corresponding to each node.
[0220] In some embodiments, the different multiple nodes include a first leaf node, a second leaf node and an intermediate node; if the different multiple nodes belong to different node levels, the third determination module 4553 is also used to determine, for the first leaf node, the first influence value of the first leaf node and the average of the influence values of the adjacent leaf nodes of the first leaf node, and use the average as the third influence value of the fifth influence factor corresponding to the first leaf node; for the second leaf node, determine the second influence value of the second leaf node and the average of the influence values of the adjacent leaf nodes of the second leaf node, and use the average as the third influence value of the fifth influence factor corresponding to the second leaf node; for the intermediate node, respectively determine the influence values of the influence factors corresponding to the two child nodes of the intermediate node, and determine the difference in the influence values of the influence factors corresponding to the two child nodes, and determine the third influence value of the fifth influence factor corresponding to the intermediate node based on the difference.
[0221] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device 400 to perform the service data processing method described above in the embodiment of the present application.
[0222] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the method for processing business data provided by the embodiment of the present application, for example, Figure 3 The method of processing business data is shown.
[0223] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0224] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0225] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0226] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0227] To sum up, through the above method, when performing data analysis on the first business, the business parameters reflecting the operating status of the first business are determined, and the first decision tree corresponding to the business parameters is determined. When the parameter value of the business parameter changes from the first parameter value to the second parameter value, a first decision path matching the first parameter value and a second decision path matching the second parameter value can be determined in the first decision tree, thereby determining the first impact value of the first influencing factor in the first decision path on the business parameter and the second impact value of the second influencing factor in the second decision path on the business parameter. Finally, based on the first impact value and the second impact value, the parameter value that makes the business parameter change can be determined from the first influencing factor and the second influencing factor. The third influencing factor whose value changes. This method can use the decision tree to automatically capture the nonlinear relationship and interaction between the business parameters and the influencing factors. In this way, when analyzing the changes in the business parameters of the first business, there is no need to preset the precise mathematical expression between the business parameters and the influencing factors as in the related technology. When the business parameters change, the third influencing factor that causes the business parameters to change can be accurately captured based on the first decision tree, so that the changes in the business parameters can be analyzed according to the third influencing factor. This method of processing business data can be applicable to different business scenarios and different data characteristics, thereby improving the flexibility and versatility of the data processing method in practical applications.
[0228] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A method for processing business data, characterized in that: The method comprises: Determining a service parameter for reflecting an operating state of a first service and a first decision tree corresponding to the service parameter, wherein the first decision tree includes multiple decision paths; When the parameter value of the service parameter changes from a first parameter value to a second parameter value, determining, based on a first decision path in the first decision tree that matches the first parameter value, a first impact value of a first influencing factor in the first decision path on the service parameter, and determining, based on a second decision path in the first decision tree that matches the second parameter value, a second impact value of a second influencing factor in the second decision path on the service parameter; Based on the first impact value and the second impact value, a third impact factor for causing a change in the parameter value of the service parameter is determined from the first impact factor and the second impact factor.
2. The method according to claim 1, characterized in that The determining of the first decision tree corresponding to the service parameter includes: Constructing a data set based on a plurality of fourth influencing factors preset for constructing the first decision tree, the data set including a first value of each of the fourth influencing factors at different time points, and a third parameter value of the business parameter at the different time points; Based on the data set, performing relationship fitting on the business parameter and the plurality of fourth influencing factors to obtain a second decision tree; Determine a coefficient of determination of the second decision tree, and adjust the structure of the second decision tree based on the coefficient of determination to obtain the first decision tree.
3. The method according to claim 2, characterized in that Determining a coefficient of determination of the second decision tree includes: Dividing the data set into at least two sub-sample sets; For each of the subsample sets, determining a sub-determination coefficient of the second decision tree based on the subsample set and the second decision tree; The adjusting the structure of the second decision tree based on the determination coefficient to obtain the first decision tree includes: If there is at least one sub-determination coefficient that is less than the first threshold, the second decision tree is iteratively fitted based on the data set until each sub-determination coefficient is greater than or equal to the first threshold, and the decision tree obtained by stopping the fitting is used as the first decision tree.
4. The method according to claim 1, wherein The first decision tree has the business parameter as a root node, and the first decision tree includes at least one node level, and each node level includes at least one influencing factor; Before determining, based on the first decision path in the first decision tree that matches the first parameter value, a first impact value of a first influencing factor in the first decision path on the service parameter, the method further includes: Determining the nodes in each node level in order from top to bottom based on the first parameter value and the segmentation conditions of the influencing factors in each node level; The nodes in the determined node levels are connected to obtain a first decision path that matches the first parameter value.
5. The method according to claim 4, characterized in that The determining, based on a first decision path in the first decision tree that matches the first parameter value, a first impact value of a first influencing factor in the first decision path on the service parameter includes: For each node level, determining a second influence value of the first influence factor in the node level based on the nodes in the node level that are located in the first decision path and the first influence factors of the nodes in the first decision path; The second impact value of the first impact factor in each of the node levels is added together to obtain a first impact value of the first impact factor in the first decision path on the business parameter.
6. The method according to claim 1, wherein The determining, based on the first impact value and the second impact value, a third impact factor causing a change in the parameter value of the service parameter from the first impact factor and the second impact factor, includes: When the first impact value is different from the second impact value, based on the difference between the first decision path and the second decision path, determining a fifth impact factor that affects a change in the parameter value of the service parameter from among the first impact factor and the second impact factor; Determining a third impact value of each of the fifth impact factors based on the first impact value and the second impact value; The fifth impact factor whose third impact value is greater than the second threshold is determined as the third impact factor of the change in the parameter value of the service parameter.
7. The method according to claim 6, characterized in that The determining, based on the difference between the first decision path and the second decision path, a fifth influencing factor affecting the change in the parameter value of the service parameter from among the first influencing factor and the second influencing factor includes: determining a plurality of different nodes in the first decision path and the second decision path, and parent nodes of the plurality of nodes; The impact factors corresponding to the multiple nodes and the impact factor corresponding to the parent node are used as the fifth impact factor.
8. The method according to claim 7, characterized in that The determining, based on the first impact value and the second impact value, a third impact value of each of the fifth impact factors includes: If the different multiple nodes belong to the same node level, determining a difference between the second influence value and the first influence value, and an average of the second influence value and the first influence value; The difference is used as the third influence value of the fifth influence factor corresponding to the parent node, and the average value is used as the third influence value of the fifth influence factor corresponding to each of the nodes.
9. The method according to claim 7, characterized in that The different multiple nodes include a first leaf node, a second leaf node and an intermediate node; If the different nodes belong to different node levels, determining the third influence value of each of the fifth influence factors based on the first influence value and the second influence value includes: For the first leaf node, determine an average of a first influence value of the first leaf node and influence values of adjacent leaf nodes of the first leaf node, and use the average as a third influence value of the fifth influence factor corresponding to the first leaf node; For the second leaf node, determine an average of the second influence value of the second leaf node and the influence values of the adjacent leaf nodes of the second leaf node, and use the average as the third influence value of the fifth influence factor corresponding to the second leaf node; For the intermediate node, the influence values of the influence factors corresponding to the two child nodes of the intermediate node are determined respectively, and the difference in the influence values of the influence factors corresponding to the two child nodes is determined; and the third influence value of the fifth influence factor corresponding to the intermediate node is determined based on the difference.
10. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the method for processing business data according to any one of claims 1 to 9 when executing the computer executable instructions or computer program stored in the memory.
11. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the method for processing business data according to any one of claims 1 to 9 is implemented.