Transaction data processing method and device, storage medium and program product
Through the Bayesian model and random search variable selection strategy, the key factors for the target object to select the financial institution to process the transaction are screened out, which solves the low efficiency problem in the existing technology and realizes efficient factor determination.
Patent Information
- Application Number
- CN202510872571.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-17
AI Technical Summary
The existing technology is inefficient in determining the factors influencing the target object's selection of a financial institution to process a transaction, and the calculation efficiency is particularly low when there are many variables.
A random search variable selection strategy based on the Bayesian model is adopted. By establishing a linear regression model, using Gibbs sampling and MCMC methods, the target factor set is determined and the factors that have a significant impact on the target object's selection of financial institutions to process transactions are screened out.
The efficiency of determining the factors that influence the target object's selection of financial institutions to process transactions is improved, the inefficiency of calculating all possible combinations is avoided, and the key factors are efficiently screened.
Smart Images

Figure CN120807146A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of financial technology or other related technical fields, in particular, to a transaction data processing method and device, a storage medium and a program product. BACKGROUND
[0002] When looking for the influencing factors of the customers' selection of the financial institutions for collecting, a model is generally established, and then the independent variables in the model are selected to determine the influencing factors of the customers' selection of the financial institutions for collecting. In related technologies, stepwise regression method can be used to select the variables to determine the influencing factors of the customers' selection of the financial institutions for collecting. However, this method can only be used for variables without correlation. In related technologies, all possible regression equation method can also be used to calculate all regression equations corresponding to all possible combinations of the independent variables, and then the optimal regression equation is determined by comparing some information criteria, so as to determine the influencing factors of the customers' selection of the financial institutions for collecting according to the independent variables in the optimal regression equation. However, in the case of many variables, the calculation will be quite laborious and inefficient.
[0003] At present, there is no effective solution to the problem of low efficiency in determining the influencing factors of the target objects' selection of the financial institutions for processing transactions in related technologies. SUMMARY
[0004] The main purpose of the present application is to provide a transaction data processing method and device, a storage medium and a program product to solve the problem of low efficiency in determining the influencing factors of the target objects' selection of the financial institutions for processing transactions in related technologies.
[0005] In order to achieve the above purpose, according to one aspect of the present application, a transaction data processing method is provided. The method comprises: obtaining transaction data of N target objects to obtain a transaction data set, wherein N is a positive integer; establishing a target regression model based on the transaction data set, wherein the type of the target regression model includes a linear regression model; determining a target factor set based on a random search variable selection strategy and the target regression model, wherein the target factor set includes the influencing factors of the target objects' selection of the financial institutions for processing transactions.
[0006] Further, the random search variable selection strategy includes a Bayesian model-based variable selection strategy. Determining the target factor set based on the random search variable selection strategy and the target regression model comprises: fitting the parameters of the target regression model based on the Bayesian model to obtain a fitting result; selecting the variables of the target regression model based on the fitting result to obtain a selection result; and determining the target factor set based on the variables in the selection result.
[0007] Further, the parameters of the target regression model include coefficients of independent variables in the target regression model, the parameters of the target regression model are fitted by using the Bayesian model to obtain a fitting result, including: extracting the coefficients of the independent variables in the target regression model to obtain the parameters of the target regression model; determining initial parameters of the Bayesian model based on the parameters of the target regression model, wherein the initial parameters include the parameters of the target regression model, an indicator variable and a variance, the indicator variable is used to select variables of the target regression model; performing M times of iterative updates on the initial parameters of the Bayesian model to obtain the fitting result, wherein M is a positive integer.
[0008] Further, the M times of iterative updates on the initial parameters of the Bayesian model to obtain the fitting result, including: determining N prior distributions of the Bayesian model, wherein the N prior distributions include a prior distribution of the parameters of the target regression model, a prior distribution of the indicator variable and a prior distribution of the variance; performing M times of iterative updates on the initial parameters of the Bayesian model based on the N prior distributions of the Bayesian model, and determining a posterior distribution of the indicator variable; determining the fitting result based on the posterior distribution of the indicator variable.
[0009] Further, the prior distribution of the parameters of the target regression model includes a mixture normal distribution, the prior distribution of the indicator variable includes a binomial distribution, and the prior distribution of the variance includes an inverse gamma distribution.
[0010] Further, the indicator variable includes S variables, each variable is associated with a coefficient of an independent variable in the target regression model, S is a positive integer, and the fitting result includes an iteratively updated indicator variable, the variables of the target regression model are selected based on the fitting result to obtain a selection result, including: in a case where a certain variable in the iteratively updated indicator variable is a first preset value, performing exclusion processing on an independent variable associated with the certain variable in the target regression model to obtain a first processing result; in a case where a certain variable in the iteratively updated indicator variable is a second preset value, performing retention processing on an independent variable associated with the certain variable in the target regression model to obtain a second processing result; and determining the selection result based on the first processing result and the second processing result.
[0011] Further, the target regression model is established based on the transaction dataset, including: establishing an initial regression model, wherein the initial regression model includes S independent variables and to-be-solved coefficients of each independent variable; solving the to-be-solved coefficients of each independent variable in the initial regression model based on the transaction data to obtain a solving result; and determining the target regression model based on the solving result.
[0012] To achieve the above object, according to another aspect of the present application, a transaction data processing device is provided. The device includes: an acquisition unit configured to acquire transaction data of N target objects to obtain a transaction dataset, wherein N is a positive integer; an establishment unit configured to establish a target regression model based on the transaction dataset, wherein the type of the target regression model includes a linear regression model; and a determination unit configured to determine a target factor set based on a random search variable selection strategy and the target regression model, wherein the target factor set includes an impact factor of a financial institution that selects a target object to process a transaction.
[0013] Further, the random search variable selection strategy includes a variable selection strategy based on a Bayesian model, the determination unit includes: a fitting subunit configured to fit parameters of the target regression model by using the Bayesian model to obtain a fitting result; a selection subunit configured to select variables of the target regression model based on the fitting result to obtain a selection result; and a first determination subunit configured to determine the target factor set based on the variables in the selection result.
[0014] Further, the parameters of the target regression model include coefficients of independent variables in the target regression model, the fitting subunit includes: an extraction module configured to extract the coefficients of the independent variables in the target regression model to obtain the parameters of the target regression model; a first determination module configured to determine initial parameters of the Bayesian model based on the parameters of the target regression model, wherein the initial parameters include the parameters of the target regression model, an indicator variable, and a variance, and the indicator variable is used to select variables of the target regression model; and an update module configured to perform M times of iterative updates on the initial parameters of the Bayesian model to obtain the fitting result, wherein M is a positive integer.
[0015] Further, the update module includes: a first determination subunit configured to determine N prior distributions of the Bayesian model, wherein the N prior distributions include a prior distribution of the parameters of the target regression model, a prior distribution of the indicator variable, and a prior distribution of the variance; a processing submodule configured to perform M times of iterative updates on the initial parameters of the Bayesian model based on the N prior distributions of the Bayesian model, and determine a posterior distribution of the indicator variable; and a second determination submodule configured to determine the fitting result based on the posterior distribution of the indicator variable.
[0016] Further, the distribution type of the prior distribution of the parameters of the target regression model comprises a mixture normal distribution, the distribution type of the prior distribution of the latent variable comprises a binomial distribution, and the distribution type of the prior distribution of the variance comprises an inverse gamma distribution.
[0017] Further, the latent variable comprises S variables, each variable is associated with a coefficient of one independent variable in the target regression model, S is a positive integer, the fitting result comprises the iteratively updated latent variable, the selection unit comprises: a removing module configured to, in a case where one variable in the iteratively updated latent variable is a first preset value, remove the independent variable associated with the variable in the target regression model to obtain a first processing result; a processing module configured to, in a case where one variable in the iteratively updated latent variable is a second preset value, retain the independent variable associated with the variable in the target regression model to obtain a second processing result; and a second determination module configured to determine the selection result based on the first processing result and the second processing result.
[0018] Further, the establishing unit comprises: an establishing subunit configured to establish an initial regression model, wherein the initial regression model comprises S independent variables and to-be-solved coefficients of each independent variable; a solving subunit configured to solve the to-be-solved coefficients of each independent variable in the initial regression model based on the transaction data to obtain a solving result; and a second determination subunit configured to determine the target regression model based on the solving result.
[0019] According to another aspect of the present application, a computer readable storage medium is provided, which comprises a stored executable program, wherein the computer readable storage medium controls a device where the computer readable storage medium is located to execute the transaction data processing method when the executable program is run.
[0020] According to another aspect of the present application, an electronic device is provided, which comprises: a memory storing an executable program; and a processor configured to run the program, wherein the program is configured to execute the transaction data processing method when the program is run.
[0021] According to another aspect of the present application, a computer program product is provided, which comprises computer instructions configured to implement the steps of the transaction data processing method when executed by a processor.
[0022] In the embodiment of the present application, the transaction data of N target objects is acquired to obtain a transaction data set, wherein N is a positive integer; a target regression model is established based on the transaction data set, wherein the type of the target regression model includes a linear regression model; based on a random search variable selection strategy and the target regression model, a target factor set is determined, wherein the target factor set includes an impact factor of a financial institution that processes transactions of the target objects, thereby solving the technical problem of low efficiency in determining the impact factor of the financial institution that processes transactions of the target objects in the related art. In the present application, the impact factor of the financial institution that processes transactions of the target objects is determined through the random search variable selection strategy, avoiding the situation of low efficiency in determining the impact factor of the financial institution that processes transactions of the target objects in the related art according to all possible combinations of variables, thereby achieving the technical effect of improving the efficiency of determining the impact factor of the financial institution that processes transactions of the target objects. BRIEF DESCRIPTION OF DRAWINGS
[0023] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and are used to interpret the illustrative embodiments of the present application and their descriptions, and do not constitute improper limitations to the present application. In the drawings:
[0024] Figure 1 A hardware structure block diagram of a computer terminal for implementing the processing method of transaction data is shown;
[0025] Figure 2 A flowchart of the processing method of transaction data according to the embodiment of the present application is shown;
[0026] Figure 3 A flowchart of the posteriori estimation of γ[j] according to the embodiment of the present application is shown;
[0027] Figure 4 A schematic diagram of the processing device of transaction data according to the embodiment of the present application is shown;
[0028] Figure 5 A structure block diagram of an electronic device according to the embodiment of the present application is shown. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] It should be noted that the transaction data processing method and device in the present application can be used in the field of financial technology to process the data of the acquirer service of the financial institution, and can also be used in any field other than the field of financial technology to process the data of the acquirer service of the financial institution. The application field of the transaction data processing method and device in the present application is not limited.
[0032] First, some of the nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:
[0033] The basic idea of the MCMC method is to obtain samples of π(θ) by establishing a Markov chain with a stationary distribution π(θ). Based on these samples, various statistical inferences can be made.
[0034] Gibbs sampling is a special MCMC algorithm. The basic idea is that when sampling a high-dimensional population or a complex population, a Markov chain {θ (j)} is constructed by using the conditional distribution family of distribution π, so that it has π as an invariant distribution.
[0035] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, and transaction data, etc.) involved in the present application are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal. For example, the system and the interface between the related users or institutions provide the user with a corresponding operation portal for the user to choose to agree or refuse the automatic decision result; if the user chooses to refuse, the expert decision process is entered.
[0036] It should be noted that the user can view the transaction data use purpose in real time through the authorization interface, and has the right to withdraw authorization or delete data at any time. After withdrawing authorization, the relevant data processing will be terminated within 24 hours.
[0037] The application can be applied to various software products, control systems, and client terminals (including but not limited to mobile clients, PCs, etc.) of various financial institutions. Taking the software product as an example, through the software product installed on the mobile client, the business content of the financial institution (including but not limited to transfer, financial management, fund, payment, account inquiry, advertising, recommendation, etc.) can be realized.
[0038] Embodiment one
[0039] According to the embodiment of the application, a method embodiment of a transaction data processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0040] The method embodiment provided by the embodiment of the application can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the transaction data processing method is shown. As shown in Figure 1 The computer terminal 10 (or mobile device) can include one or more processors 102 (the processor 102 can include but is not limited to a microprocessor MCU or a programmable logic device FPGA processing device), a memory 104 for storing data, and a transmission device 106 for communication function. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can also include more or less components than those shown in Figure 1 , or have a different configuration than Figure 1 .
[0041] It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be referred to herein generally as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any one of the other elements of the computer terminal 10 (or mobile device). As referred to in the embodiments herein, the data processing circuitry functions as a processor to control, for example, the selection of the variable resistance terminal path connected to the interface.
[0042] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the transaction data processing method of the embodiments herein. The processor 102 can execute various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e. implement the transaction data processing method described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 can further include a memory disposed remotely with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0043] The transmission device 106 is configured to receive or send data via a network. Examples of the network include, but are not limited to, a wireless network provided by a communication service provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.
[0044] The display can be, for example, a touch screen type liquid crystal display (LCD) that enables a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0045] In the above-described operating environment, the embodiments herein provide a transaction data processing method as shown in Figure 2 Figure 2 is a flowchart of the transaction data processing method according to the first embodiment herein.
[0046] In step S201, transaction data of N target objects is obtained to obtain a transaction data set, where N is a positive integer.
[0047] The transaction data can include transaction data authorized by the target object, and the transaction data can include transaction data of the target object for which the financial institution provides a collection service. For example, the transaction data of the target object can include an average transaction rate X2 of the target object at the financial institution A, an average transaction rate X3 of the target object at the financial institution B, an average transaction rate X4 of the target object at the financial institution C, and an average transaction rate X5 of the target object at other institutions.
[0048] In step S202, a target regression model is established based on the transaction data set, and the type of the target regression model includes a linear regression model.
[0049] In this embodiment, an initial linear regression model in which y i and independent variables x 1i , x 2i , …, x pi have a linear relationship can be established as follows:
[0050] y i = β1x 1i + β2x 2i + … + β p x pi + ε i , i = 1, 2, …, n
[0051] After the initial linear regression model is established, β1, β2, …, β p can be solved based on the transaction data to obtain the values of β1, β2, …, β p , and thus the target regression model can be obtained.
[0052] In step S203, a target factor set is determined based on a random search variable selection strategy and the target regression model, and the target factor set includes an influencing factor of the financial institution that processes the transaction selected by the target object.
[0053] In this embodiment, the random search variable selection strategy is used, and the random search variable selection strategy is applied to the linear regression model (the target regression model) to obtain the conditional posterior distribution of the model, to convert the model selection problem into a calculation problem of the posterior probability, and to find important factors of the financial institution that affects the target selection collection.
[0054] The random search variable selection is used to select a subset of variables in the linear regression model. The basic idea is to establish a hierarchical Bayesian normal mixture model, and the indicator variable of the Bayesian normal mixture model is used to identify and determine whether to select the item (independent variable). The sub-model with the highest posterior probability (and the sub-model of the target regression model, the independent variables in the sub-model are a subset of the independent variables in the target regression model) is the optimal model. The random search variable selection is to use Gibbs sampling to sample from the posterior distribution to find the sub-model with the highest probability.
[0055] Through the above steps, in this embodiment, the influencing factors of the financial institutions affecting the target object selection processing transaction are determined by the random search variable selection strategy, which avoids the low efficiency of determining the influencing factors of the financial institutions affecting the target object selection processing transaction according to all possible combinations of variables in the related art, thereby realizing the technical effect of improving the efficiency of determining the influencing factors of the financial institutions affecting the target object selection processing transaction. Further, the technical problem of low efficiency in determining the influencing factors of the financial institutions affecting the target object selection processing transaction in the related art is solved.
[0056] For example, step (one): collect the demand of the merchant for the acquirer service of financial institution A:
[0057] An example of collecting the demand of the merchant (corresponding to the target object) for the acquirer service of financial institution A in a historical period is to study the selection factors of the merchant for the charging service institution. It is necessary to find the factors that have a significant impact on the demand Y of the financial institution acquirer service for each merchant in the data set (corresponding to the transaction data). These factors can be the total amount of each merchant transaction X1, the average rate of each transaction of financial institution A X2, the average rate of each transaction of financial institution B X3, the average rate of each transaction of financial institution C X4, and the average rate of each transaction of other institutions X5.
[0058] Step (two): establish a Bayesian model of linear random search variable selection:
[0059] 1. Establish a linear regression model.
[0060] Assume that there is a linear relationship between the random variable y and the independent variables x1, x2, …, xn: p
[0061] y i = β1x 1i + β2x 2i + … + β p x pi + ε i , i = 1, 2, …, n
[0062] 2. Determine the prior of the parameter β of the Bayesian model, the prior of the hyperparameter γ (indicator variable), and the parameter σ2 the prior, the posterior distribution of the three parameters is calculated, and the posterior model probability of the Bayesian model is obtained.
[0063] 3. The random search variable selection method generates an MCMC sample {γ l , l = 1, 2, …, L} of γ using Gibbs sampling, and screens out the items with γ j = 1, which are the items of the linear regression model. The Gibbs sampling steps are as follows:
[0064] ① Give the starting point (β 0 , σ 2 ) 0 , γ 0 );
[0065] ② Sample from β l ~ N p (β*, V*), where is the maximum likelihood estimate, X represents the matrix of independent variables x1, x2, …, x p , and D = (n, Y, X) is the observed data, and D γ is the observed data corresponding to the indicator variable;
[0066] ③ Sample from , and is a hyperparameter;
[0067] ④ Sample from .
[0068] Step (three): According to the MCMC results of the parameters β in the linear model, remove the insignificant independent variables in the linear model, and use the software package for Bayesian statistical analysis to perform Bayesian estimation on the optimal model again to obtain the final model of the demand of merchants (corresponding to the target object) for the financial institution's service, and find the most important factors affecting the selection of financial institutions' service by merchants.
[0069] 1. Define the parameter space: clearly define the parameters to be adjusted and their value range (e.g., the number of iterations);
[0070] 2. Set the random sampling strategy: determine how to randomly sample in the defined parameter space. Uniform distribution, log-uniform distribution, etc. can be used to adapt to the characteristics of different types of parameters.
[0071] 3. Perform multiple random sampling and model evaluation: randomly sample multiple parameter combinations, and use these parameters to train the model (e.g., posterior model), and then evaluate the performance of the model in the validation set or cross-validation.
[0072] 4. Record and compare results: Record the parameter combination and its corresponding model performance for each sampling, for subsequent comparison and analysis.
[0073] 5. Select the optimal parameters: According to the recorded results, select the parameter combination that optimizes the linear model performance as the final parameter tuning result.
[0074] For example, take the data of merchants' demand for financial institution A's collection service as an example.
[0075] Collect an example of merchants' (corresponding to target objects) demand for financial institution A's collection service in a historical period, aiming to study the selection factors of merchants for the charging service institution. In the data set (corresponding to transaction data), as shown in Table 1, find out the factors that have a significant impact on the financial institution's collection service demand Y for each merchant, which can be the total transaction amount X1 of each merchant, the average transaction rate of financial institution A X2, the average transaction rate of financial institution B X3, the average transaction rate of financial institution C X4, and the average transaction rate of other institutions X5.
[0076] Table 1
[0077]
[0078]
[0079] Where Y = merchants' demand for financial institution A's collection service; X1 = total transaction amount of each merchant, ten thousand yuan; X2 = average transaction rate of financial institution A per transaction, ten thousandth; X3 = average transaction rate of financial institution B per transaction, ten thousandth; X4 = average transaction rate of financial institution C per transaction, ten thousandth; X5 = average transaction rate of other institutions per transaction, ten thousandth.
[0080] Assume that y i and independent variables x 1i ,x 2i ,…,x pi There is a linear relationship between them:
[0081] y i = β1x 1i + β2x 2i + … + β p x pi + ε i , i = 1, 2, …, n
[0082] Where ε i is independent and identically distributed in N(0, σ 2 ).
[0083] The random search variable method is used to select the model, which can give the optimal subset of the model at one time. First, the standard deviation of the least square estimate of the coefficient β under the full model can be calculated The results calculated by the related software are 0.004962, 0.174400, 0.089544, 0.070644, and 0.098381, respectively. Therefore, τ j = 0.09 and j = 1, …, 5 can be selected. The least square estimate of the variance σ 2 Therefore, the initial value (σ 2 ) 0 = 4 can be selected. In addition, the prior distribution of the indicator variable γ can be specified as a uniform distribution, i.e., π(γ) = 2 -5 . Other hyperparameters are set as follows: c1 = c2 = … = c8 = 10, R is a unit matrix I, the hyperparameters υ = 3 and λ = (σ 2 ) 0 / 3 of the inverse Gamma distribution.
[0084] Two Markov chains are generated by using the software package for performing Bayesian statistical analysis, and the data obtained are iterated for 20,000 times, and the first 10,000 iterations can be discarded, and both of the two chains are convergent. Table 2 gives the posterior model probabilities of the models selected by the random search selection method. The settings of c j and τ j are important in the random search selection method, because they are related to the distribution of the linear regression coefficient, and the setting of τ j is based on However, the setting of c j is based on different problems. Table 2 shows that the different settings of c j in this example do not affect the final result. The model with the largest posterior probability is the third model, and the form of the model is {01000} after being converted into binary numbers, that is, only the X2 term is left.
[0085] Table 2
[0086]
[0087] The posterior mean values of γ[j], j = 1, 2, …, 5 are γ[1] = 0.05035, γ[2] = 0.8186, γ[3] = 0.0936, γ[4] = 0.1938, and γ[5] = 0.3636, Figure 3 The flowchart of the posterior estimation of γ[j] is shown in Figure 3 It can be seen that only the value of γ[2] is closer to 1, which once again indicates that the linear regression model should only include the X2 term.
[0088] Table 3 gives the MCMC results of parameter β in the full model. It is seen that the random variable search method can accurately find the model according to the value of γ. In order to obtain better model estimates, the insignificant independent variables are removed, and the optimal model is subjected to Bayesian estimation by using the software package for performing Bayesian statistical analysis, and the model of the demand of the merchant for the acquirer service of the financial institution A is obtained as follows:
[0089] Y = 0.813X2
[0090] Table 3
[0091] Parameter Mean Standard deviation MCerror 2.5% Median 97.5% [Alpha]1 -0.01388 0.009513 6.349E-4 -0.03569 -0.0129 0.001886 <![CDATA[β2]]> 0.5018 0.2655 0.01783 -0.006646 0.546 0.9229 [Alpha]3 -0.1265 0.09231 0.005598 -0.1844 -0.01552 0.1802 [Alpha]4 0.09939 0.1126 0.007568 -0.0801 0.08291 0.4062 [Alpha]5 0.1691 0.1574 0.0107 -0.05728 0.1312 0.5154
[0092] Where MCerror (Monte Carlo Error) represents the estimation error of the MCMC sample sequence for the posterior mean of the parameter. In this embodiment, the traditional stepwise regression method can also be used to model the data set, and the coefficients of X1, X3, X4 and X5 are not significant, only the coefficient of X2 is significant, the P value of t test is less than 0.05, and the corresponding coefficient is 0.813. F test is performed, the equation is significant, R 2 = 0.986, where the models obtained by the two methods are the same.
[0093] In this embodiment, the acquirer service of the financial institution A, the acquirer service of the financial institution C, the acquirer service of the financial institution D and the acquirer service of other institutions are substitutes for each other, and the variables X1, X3, X4 and X5 are correlated. However, through Bayesian model selection, it can be seen that the consumption of the acquirer service of the financial institution A of the merchant is only positively correlated with the rate of the acquirer service of the financial institution A, and the relationship with its substitutes and the total transaction amount of the merchant is not significant.
[0094] Although the random search selection method does not need to calculate the information amount of all possible sub-models, it searches for the optimal regression equation from all possible sub-models, and the judgment basis is also relatively objective according to the posterior model probability. The software package for performing Bayesian statistical analysis is used to realize the variable selection problem of the linear model without additional programming sampling, and the variable selection problem of the linear model can be well and conveniently solved.
[0095] Optionally, in the transaction data processing method provided in the embodiment of the application, the random search variable selection strategy includes a Bayesian model-based variable selection strategy, and the target factor set is determined based on the random search variable selection strategy and the target regression model, including: fitting parameters of the target regression model by using a Bayesian model to obtain a fitting result; selecting variables of the target regression model based on the fitting result to obtain a selection result; and determining the target factor set based on the variables in the selection result.
[0096] The parameters of the target regression model include coefficients of independent variables in the target regression model. In this embodiment, a Bayesian model can be introduced to set a prior distribution based on the parameters (such as regression coefficients) of the target regression model, which can reflect the preliminary setting of the parameter values. Then, the prior distribution of the parameters can be updated to obtain a posterior distribution by using the Bayesian rule. This process can be referred to as fitting of model parameters, and can be implemented by means of an MCMC (Markov Chain Monte Carlo) method, such as Gibbs sampling.
[0097] After obtaining the posterior distribution of the parameters, the posterior probability of each parameter can be checked, especially those close to zero. Because if the posterior probability of a parameter is close to zero, the corresponding independent variable can have little contribution to the prediction of the model. Through the evaluation of the posterior probability of the parameters, it can be determined which independent variables should be retained in the target regression model and which should be excluded. In essence, variable selection is performed, and the goal is to obtain an optimal subset model, that is, to include variables that have a significant contribution to the prediction of the model.
[0098] After the variable selection step, a subset including the most relevant prediction variables can be obtained. This set of variables can also be referred to as a target factor set, which is the most important component in the target regression model and can best explain or predict the changes of the corresponding variables. The establishment of the target factor set described above means that the key factors affecting the selection of acquirers by merchants have been found.
[0099] The random search strategy is used to select variables based on Bayesian theory. Compared with the traditional method of adding or deleting variables one by one (such as stepwise regression), the random search variable selection strategy is more effective in high-dimensional data, because it can avoid the problem of combinatorial explosion, and at the same time, by using MCMC sampling to explore the parameter space, it can find the optimal variable combination, not only identify key variables, but also estimate the importance of these variables in the model, that is, by judging the size of the posterior probability, the efficiency of determining the factors affecting the selection of financial institutions by target objects is improved.
[0100] Optionally, in the method for processing transaction data provided in the embodiments of the present application, the parameters of the target regression model include coefficients of independent variables in the target regression model, the Bayesian model is used to fit the parameters of the target regression model to obtain a fitting result, including: extracting the coefficients of the independent variables in the target regression model to obtain the parameters of the target regression model; determining the initial parameters of the Bayesian model based on the parameters of the target regression model, wherein the initial parameters include the parameters of the target regression model, an indicator variable, and a variance, and the indicator variable is used to select variables of the target regression model; performing M times of iterative updates on the initial parameters of the Bayesian model to obtain the fitting result, wherein M is a positive integer.
[0101] For example, in the Bayesian model, in order to select variables for the target regression model, the indicator variables (γ1,γ 2, ...,γ k ), these variables take values of 0 or 1. γ i =1 means that the ith independent variable should be included in the model, and γ i = 0 means that the i-th independent variable is not included and needs to be eliminated, that is, the corresponding βi is set to 0. The variance (σ 2 ) is another important parameter that reflects the uncertainty of the model prediction. In the Bayesian framework, the variance is usually assigned a prior distribution (such as the inverse gamma distribution), and the prior knowledge about the variance can be integrated into the model. The parameters of the Bayesian model are iteratively updated through MCMC (Markov Chain Monte Carlo) methods, such as Gibbs sampling. In each iteration, the posterior distribution of each parameter can be updated based on the current parameter estimate and the data. In each iteration, samples can be drawn from the posterior distribution of the current parameter (for example, when sampling β, γ, and σ 2 After sampling, these new parameter values can be used to fit the model (such as calculating predicted values or goodness of fit) and update the posterior distribution of the parameters according to Bayesian theory. Each iteration is to make the parameter estimate closer to the true posterior distribution. After M iterations, the parameter estimates (including β, γ and σ) are obtained. 2 ), which is the fitting result of the Bayesian model. The posterior distribution of these parameters provides a quantification of the uncertainty of the model parameters, while the posterior distribution of the indicator variables reflects the probability of each independent variable being included in the model.
[0102] Determining the initial parameters of a Bayesian model based on the parameters of the target regression model is a method that combines statistical estimation with Bayesian statistical thinking. The introduction of indicator variables makes variable selection possible, while iteratively updating the parameters of the Bayesian model through MCMC not only yields more accurate parameter estimates but also quantifies model uncertainty, which is extremely valuable for subsequent decision analysis and interpretation of the predictive model. M iterations of updates ensure the stability and convergence of the parameter estimates.
[0103] In this embodiment, if the regression model is:
[0104] y i =β1x 1i +β2x 2i +…+β p x pi +ε i ,i=1,2,…,n
[0105] Among them, β1, β2, ..., β pThe coefficients of the independent variables in the target regression model can be called.
[0106] where ε i are independent and identically distributed N(0,σ 2 ), the model can be expressed in matrix form as:
[0107] Y = Xβ + ε, ε ~ N n (0,σ 2 I n )
[0108] where
[0109] where β = (β1, …, β p )' and σ 2 > 0 are unknown parameters. In the model Y = Xβ + ε, selecting a subset of the independent variables is equivalent to setting the coefficients β j (j = 1, 2, …, p) corresponding to the independent variables not selected to be 0. Define D = (n, Y, X) as the observed data.
[0110] In this embodiment, the Bayesian principle can be used according to the formula Y = Xβ + ε. Obviously, Y ~ N n (Xβ, σ 2 I n ), therefore, the likelihood function is:
[0111]
[0112] where
[0113] When the likelihood function and the prior distribution of each unknown parameter are given, the conditional posterior distribution density function of each variable is calculated as follows.
[0114] (1) The posterior distribution of β:
[0115]
[0116] where A r = (σ -2 X'X + (D γ RD γ ) -1 ) -1 .
[0117] (2) The posterior distribution of σ 2 :
[0118]
[0119] (3) The posterior distribution of γ:
[0120] Let
[0121] a j = f(Y | β, σ 2 ) π(β | γ (-j) , γ j = 1) π(σ 2 | γ (-j) , γ j = 1) π(γ (-j) , γ j = 1)
[0122] b j = f(Y | β, σ 2 ) π(β | γ (-j) , γ j = 0) π(σ 2 | γ (-j) , γ j = 0) π(γ (-j) , γ j = 0)
[0123] So we have γ j ~ Binomial(1, P(γ j = 1 | β, σ 2 , γ (-j) , Y)).
[0124] So
[0125] The conditional posterior distribution of the indicator variable γ is obtained, i.e. the posterior model probability of the model, so that the optimal sub-model of the model Y = Xβ + ε can be obtained. The random search variable selection method is to generate an MCMC sample {γ l , l = 1, 2, …, L} of γ by Gibbs sampling, and the items with γ j = 1 are screened out, i.e. the items of the linear model. The steps of Gibbs sampling are as follows:
[0126] ① Given the starting point (β 0 , (σ 2 ) 0 , γ 0 ),
[0127] ② Sample β l ~ N p (β*, V*), where is the maximum likelihood estimate;
[0128]
[0129] ④ Sample γ from
[0130] The random search variable method avoids calculating the posterior model probabilities of all possible submodels, and can also set a maximum number of possible variables and appropriately specify different parameters to control the calculation of the random search variable method.
[0131] Optionally, in the transaction data processing method provided in the embodiments of the present application, the initial parameters of the Bayesian model are updated for M times of iterations to obtain a fitting result, including: determining N prior distributions of the Bayesian model, wherein the N prior distributions include: a prior distribution of parameters of the target regression model, a prior distribution of the indicator variable, and a prior distribution of the variance; updating the initial parameters of the Bayesian model for M times of iterations based on the N prior distributions of the Bayesian model, and determining a posterior distribution of the indicator variable; and determining the fitting result based on the posterior distribution of the indicator variable.
[0132] For example, each parameter in the Bayesian model has its corresponding prior distribution. In this context, the N prior distributions correspond to the parameters of the target regression model (i.e. the coefficients of the independent variables), the indicator variable, and the variance, respectively. The prior of the coefficient β can be a normal distribution N(μ,σ 2 ), where μ is the preset mean value of the coefficient, and σ 2 is the preset variance. If there is no prior information, μ can be set to 0 and a large σ 2 to represent extensive uncertainty. The indicator variable γ is used for variable selection. The variance σ 2 can be assigned an inverse gamma distribution.
[0133] In this embodiment, the MCMC (Markov Chain Monte Carlo) method can be used to update the Bayesian model iteratively. In each iteration, the model parameters (β, γ and σ 2 ) are re-estimated according to their posterior distributions. This process can be performed for M times, where M is a preset positive integer to ensure that the parameter estimation reaches a stable state or converges.
[0134] As the iteration proceeds, the posterior distribution of the indicator variable γ can be updated continuously. This distribution reflects the probability that each variable is included in the model given the data. At the end of the iteration, the posterior distribution of γ can be used to determine which variables should be retained in the final model.
[0135] After M iterations are completed, the fitting result can be determined based on the posterior distribution of the indicator variable. For example, the posterior probability of each indicator variable γ i can be checked, and according to the posterior probability, it can be determined whether the corresponding independent variable X i should be included in the final model.
[0136] In the embodiment, the uncertainty of the parameter estimation (by the posterior distribution) is quantified, and the variable selection is performed by the posterior distribution of the indicator variable, so that a more concise and more predictive model is constructed. The M times of iteration update ensures that the embodiment can sufficiently explore the parameter space and converge to a stable parameter estimation state. In the case of processing a large number of variables and complex data, the embodiment is particularly useful, can effectively manage the model complexity, and avoid overfitting or underfitting.
[0137] Optionally, in the method for processing transaction data provided in the embodiment of the application, the distribution type of the prior distribution of the parameters of the target regression model comprises a mixed normal distribution, the distribution type of the prior distribution of the indicator variable comprises a binomial distribution, and the distribution type of the prior distribution of the variance comprises an inverse gamma distribution.
[0138] (1) Prior of β (the prior distribution of the parameters of the target regression model).
[0139] In order to determine which terms are contained in the model, each component of the coefficient matrix β can be regarded as coming from a mixed normal distribution. The mixed distribution is composed of two normal distributions with different variances. An indicator variable γ j (j=1, 2, …, p) is introduced, which takes a value of 0 or 1. The role of the indicator variable γ j is that when the selected model does not contain X j , then γ j = 0; when the selected model contains X j , then γ j = 1. The mixed normal distribution can be expressed as
[0140]
[0141] When γ j = 0, it means that the term is not in the model, and at this time Therefore, in the hyperparameter setting, τ j > 0 and τ j should be as small as possible, so that when γ j = 0, β j will be most likely estimated as 0; when γ j = 1, it means that the term is in the model, Therefore, in the hyperparameter setting, c j > 1 and c j should be as large as possible, so that when γ j = 1, β j will be most likely estimated as non-zero, and X j is included in the finally selected model.
[0142] Further, the mixed normal distribution formula can be rewritten in matrix form, that is:
[0143] β|γ ~ N p (0, D γ RD γ )
[0144] where γ = (γ1,..., γ p ), R is the prior correlation matrix, R can be generally taken as R = I or D γ = diag(a1τ1,..., a p τ p ). Here a j and γ j (j = 1, 2,..., p) have the following relationship:
[0145]
[0146] (2) Prior of hyperparameter γ (i.e. prior distribution of indicator variable)
[0147] γ is a discrete random variable, which follows Binomial distribution:
[0148] P(γ j = 1) = p j and P(γ j = 0) = 1 - p j
[0149] Therefore, In particular, taking then the prior of γ can become uniform distribution,
[0150] (3) Prior of σ 2 (corresponding to the prior distribution of variance)
[0151] According to the distribution content of conjugate prior, for normal distribution, when the mean is known and the variance is unknown, the conjugate prior distribution of scale parameter σ is inverse Gamma distribution family, that is,
[0152]
[0153] where v γ , λ γ are hyperparameters. In order to simplify the calculation, υ γ = υ and λ γ = λ are generally selected.
[0154] Optionally, in the transaction data processing method provided in the embodiment of the present application, the indicative variables include: S variables, each variable is associated with the coefficient of one of the independent variables in the target regression model, S is a positive integer, and the fitting result includes: the iteratively updated indicative variables, and the variables of the target regression model are selected based on the fitting result to obtain a selection result, including: when a variable in the iteratively updated indicative variables is a first preset value, the independent variable associated with the variable in the target regression model is eliminated to obtain a first processing result; when a variable in the iteratively updated indicative variables is a second preset value, the independent variable associated with the variable in the target regression model is retained to obtain a second processing result; and the selection result is determined based on the first processing result and the second processing result.
[0155] The above-mentioned indicator variable is usually a binary vector γ=(γ1,γ2,...,γ S) , where S is the total number of independent variables, γ i Can take the value 0 or 1. i The value of determines the i-th independent variable X i Whether to be included in the target regression model:
[0156] γ i =1 (corresponding to the second preset value) indicates that Xi is included in the model, that is, X is used in the model. i The coefficient of .
[0157] γ i =0 (corresponding to the first preset value) means X i Not included in the model, that is, X in the model i The coefficient of is set to 0, which is equivalent to removing X from the model. i .
[0158] When a certain item γ in the indicator variable after iterative update i It is the first preset value (usually 0), indicating that in the posterior distribution, the probability of the i-th independent variable being selected is low. Therefore, this independent variable X can be removed from the model. i , and obtain the first processing result. This step is essentially a "hard elimination" of the variable, which means that the variable will no longer participate in the model's prediction.
[0159] When a certain item γ in the indicator variable after iterative update i The second preset value (usually 1) indicates that in the posterior distribution, the probability of the i-th independent variable being selected is higher, so you can choose to keep this independent variable X i In the model, the second processing result is obtained. Retaining the independent variables with higher probability helps to build a more accurate prediction model.
[0160] Based on the first processing result and the second processing result, a final selection result is determined, that is, which independent variables are selected to construct an optimal regression model. By eliminating unimportant variables, the explanatory power of the model can be improved, and the uncertainty of parameter estimation can be reduced. In this way, the main influencing factors affecting the selection of the target object to select a collection institution can be determined.
[0161] Optionally, in the transaction data processing method provided in the embodiments of the present application, the target regression model is established based on the transaction data set, including: establishing an initial regression model, wherein the initial regression model includes S independent variables and the to-be-solved coefficients of each independent variable; based on the transaction data, the to-be-solved coefficients of each independent variable in the initial regression model are solved to obtain a solving result; and based on the solving result, the target regression model is determined.
[0162] Firstly, an initial regression model can be established, which includes S independent variables, and each independent variable has an unknown coefficient, which can be obtained by data estimation. In a multiple linear regression model, the model can be represented as Y=β0+β1X1+β2X2+...+β S Xs+ε, where Y is the dependent variable, X1,X2,...,X s S are independent variables, β0 is the intercept term, β1,β2,...,β s S are the corresponding independent variable coefficients, and ε is the error term, which is usually assumed to be an independent and identically distributed normal random variable, used to capture the part of the change that is not explained by the independent variables in the model.
[0163] In this embodiment, the independent variable and dependent variable values in the data can be input into the model. These data can include transaction amount, transaction frequency, different bank commission rates, etc. The coefficients of each independent variable in the initial regression model are solved using statistical methods such as least squares or maximum likelihood estimation. Thus, the most suitable coefficient estimate is obtained. After the solving process is completed, the coefficient estimate of each independent variable is obtained, that is, the solving result, based on which a final target regression model can be determined, which includes independent variables that have a significant impact on the dependent variable and their corresponding coefficients.
[0164] By establishing a regression model based on the transaction data set, the influence of different factors on the selection of the financial institution by the merchant to select a collection institution can be quantified, thereby providing strategic recommendations for the financial institution and improving user satisfaction.
[0165] In this embodiment, a Bayesian model with random search variable selection is used to analyze customer selection of acquirer financial institutions. By quantifying model analysis of merchants, important factors affecting customer selection of acquirer institutions are found, and a reliable model is provided to improve the service experience of merchants and retain merchants. Random search selects parameters randomly to avoid traversing all possible parameter combinations, so that better parameters can be found in the same time, especially in the case of large parameter space, which can significantly save computing resources. Random search allows hyperparameters to have different distributions (such as uniform distribution, logarithmic distribution, etc.), while grid search usually relies on fixed values, which makes random search more flexible when dealing with complex hyperparameter space. Random search is more likely to explore unexpected good combinations, especially in high-dimensional parameter space, which increases the chances of finding good parameter combinations through random sampling. Since it is not necessary to exhaust all possible parameter combinations, random search can find better performing models faster under limited computing resources, avoiding the excellent combinations missed by grid search due to fixed step size. Random search has become an important tool for hyperparameter optimization, especially suitable for large or complex parameter space.
[0166] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown.
[0167] Embodiment Two
[0168] The embodiments of the present application further provide a transaction data processing apparatus. It should be noted that the transaction data processing apparatus of the embodiments of the present application can be used to execute the transaction data processing method provided by the embodiments of the present application. The transaction data processing apparatus provided by the embodiments of the present application is introduced as follows.
[0169] According to the embodiments of the present application, a device for implementing the above transaction data processing method is further provided, as shown in Figure 4 The device comprises an acquisition unit 41, a establishing unit 42 and a determining unit 43.
[0170] The acquisition unit 41 is configured to acquire transaction data of N target objects to obtain a transaction data set, wherein N is a positive integer.
[0171] The establishing unit 42 is configured to establish a target regression model based on the transaction data set, wherein the type of the target regression model comprises a linear regression model.
[0172] The determining unit 43 is configured to determine a target factor set based on the random search variable selection strategy and the target regression model, wherein the target factor set includes an impact factor of a financial institution that processes a transaction of a target object.
[0173] In the transaction data processing apparatus provided in the embodiments of the present application, the transaction data of N target objects can be obtained by the obtaining unit 41 to obtain a transaction data set, wherein N is a positive integer, and the target regression model is established based on the transaction data set by the establishing unit 42, wherein the type of the target regression model includes a linear regression model, and the target factor set is determined based on the random search variable selection strategy and the target regression model by the determining unit 43, wherein the target factor set includes an impact factor of a financial institution that processes a transaction of a target object. Thus, the technical problem of low efficiency in determining the impact factor of the financial institution that processes the transaction of the target object is solved. In the embodiments, the impact factor of the financial institution that processes the transaction of the target object is determined by the random search variable selection strategy, which avoids the low efficiency in determining the impact factor of the financial institution that processes the transaction of the target object in the related art, thereby achieving the technical effect of improving the efficiency of determining the impact factor of the financial institution that processes the transaction of the target object.
[0174] Optionally, in the transaction data processing apparatus provided in the embodiments of the present application, the random search variable selection strategy includes a variable selection strategy based on a Bayesian model, the determining unit includes a fitting subunit configured to fit the parameters of the target regression model by using the Bayesian model to obtain a fitting result, a selection subunit configured to select the variables of the target regression model based on the fitting result to obtain a selection result, and a first determining subunit configured to determine the target factor set based on the variables in the selection result.
[0175] Optionally, in the transaction data processing apparatus provided in the embodiments of the present application, the parameters of the target regression model include the coefficients of independent variables in the target regression model, the fitting subunit includes an extraction module configured to extract the coefficients of the independent variables in the target regression model to obtain the parameters of the target regression model, a first determining module configured to determine initial parameters of the Bayesian model based on the parameters of the target regression model, wherein the initial parameters include the parameters of the target regression model, an indicator variable, and a variance, and the indicator variable is used to select the variables of the target regression model, and an updating module configured to perform M times of iterative updating on the initial parameters of the Bayesian model to obtain the fitting result, wherein M is a positive integer.
[0176] Optionally, in the transaction data processing apparatus provided by the embodiment of the present application, the updating module comprises: a first determining subunit configured to determine N prior distributions of the Bayesian model, wherein the N prior distributions comprise a prior distribution of a parameter of the target regression model, a prior distribution of the indicator variable, and a prior distribution of the variance; a processing submodule configured to perform M times of iterative updating on initial parameters of the Bayesian model based on the N prior distributions of the Bayesian model, and determine a posterior distribution of the indicator variable; and a second determining subunit configured to determine the fitting result based on the posterior distribution of the indicator variable.
[0177] Optionally, in the transaction data processing apparatus provided by the embodiment of the present application, the distribution type of the prior distribution of the parameter of the target regression model comprises a mixture normal distribution, the distribution type of the prior distribution of the indicator variable comprises a binomial distribution, and the distribution type of the prior distribution of the variance comprises an inverse gamma distribution.
[0178] Optionally, in the transaction data processing apparatus provided by the embodiment of the present application, the indicator variable comprises S variables, each variable is associated with a coefficient of one independent variable in the target regression model, S is a positive integer, the fitting result comprises the iteratively updated indicator variable, the selecting subunit comprises: a removing module configured to, in a case where one variable in the iteratively updated indicator variable is a first preset value, perform removing processing on the independent variable associated with the variable in the target regression model to obtain a first processing result; a processing module configured to, in a case where one variable in the iteratively updated indicator variable is a second preset value, perform retaining processing on the independent variable associated with the variable in the target regression model to obtain a second processing result; and a second determining module configured to determine a selection result based on the first processing result and the second processing result.
[0179] Optionally, in the transaction data processing apparatus provided by the embodiment of the present application, the establishing unit comprises: an establishing subunit configured to establish an initial regression model, wherein the initial regression model comprises S independent variables and to-be-solved coefficients of each independent variable; a solving subunit configured to solve the to-be-solved coefficients of each independent variable in the initial regression model based on the transaction data to obtain a solving result; and a second determining subunit configured to determine the target regression model based on the solving result.
[0180] It should be noted that the obtaining unit 41, the establishing unit 42 and the determining unit 43 correspond to steps S201 to S203 in Embodiment One, and each unit has the same instances and application scenarios as the corresponding steps, but is not limited to the disclosure of Embodiment One. It should be noted that the above modules or units can be hardware components or software components stored in the memory (for example, the memory 104) and processed by one or more processors (for example, the processors 102a, 102b, …, 102n), and the above modules can also be run in the computer terminal 10 provided in Embodiment One as part of the device.
[0181] Embodiment Three
[0182] Embodiments of the present application can provide an electronic device, Figure 5 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 5 shown, the electronic device can include one or more (only one is shown in the figure) processors 502, a memory 504, a storage controller, and a peripheral interface, wherein the peripheral interface is connected with a radio frequency module, an audio module and a display. Figure 5
[0183] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned method. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0184] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: obtaining transaction data of N target objects to obtain a transaction data set, wherein N is a positive integer; establishing a target regression model based on the transaction data set, wherein the type of the target regression model includes a linear regression model; determining a target factor set based on a random search variable selection strategy and the target regression model, wherein the target factor set includes an influencing factor of a financial institution that selects and processes transactions of the target objects.
[0185] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: the random search variable selection strategy includes a variable selection strategy based on a Bayesian model; based on the random search variable selection strategy and the target regression model, a target factor set is determined, including: fitting parameters of the target regression model by using the Bayesian model to obtain a fitting result; based on the fitting result, variables of the target regression model are selected to obtain a selection result; based on the variables in the selection result, the target factor set is determined.
[0186] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: the parameters of the target regression model include coefficients of independent variables in the target regression model; fitting the parameters of the target regression model by using the Bayesian model to obtain a fitting result includes: extracting the coefficients of the independent variables in the target regression model to obtain the parameters of the target regression model; based on the parameters of the target regression model, initial parameters of the Bayesian model are determined, wherein the initial parameters include: the parameters of the target regression model, an indicator variable, and a variance, and the indicator variable is used to select variables of the target regression model; the initial parameters of the Bayesian model are updated for M times to obtain the fitting result, wherein M is a positive integer.
[0187] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: updating the initial parameters of the Bayesian model for M times to obtain the fitting result includes: determining N prior distributions of the Bayesian model, wherein the N prior distributions include: a prior distribution of the parameters of the target regression model, a prior distribution of the indicator variable, and a prior distribution of the variance; based on the N prior distributions of the Bayesian model, the initial parameters of the Bayesian model are updated for M times, and a posterior distribution of the indicator variable is determined; based on the posterior distribution of the indicator variable, the fitting result is determined.
[0188] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: the distribution type of the prior distribution of the parameters of the target regression model includes a mixed normal distribution; the distribution type of the prior distribution of the indicator variable includes a binomial distribution; and the distribution type of the prior distribution of the variance includes an inverse gamma distribution.
[0189] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: the indicator variable includes S item variables, each item variable is associated with a coefficient of one independent variable in the target regression model, S is a positive integer, the fitting result includes the iteratively updated indicator variable, selecting variables of the target regression model based on the fitting result to obtain a selection result, including: in the case that an item variable in the iteratively updated indicator variable is a first preset value, performing elimination processing on the independent variable associated with the item variable in the target regression model to obtain a first processing result; in the case that an item variable in the iteratively updated indicator variable is a second preset value, performing retention processing on the independent variable associated with the item variable in the target regression model to obtain a second processing result; determining the selection result based on the first processing result and the second processing result.
[0190] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: establishing a target regression model based on a transaction data set, including: establishing an initial regression model, wherein the initial regression model includes S independent variables and the to-be-solved coefficients of each independent variable; solving the to-be-solved coefficients of each independent variable in the initial regression model based on the transaction data to obtain a solution result; and determining the target regression model based on the solution result.
[0191] By using the embodiments of the present application, the influence factors of the financial institutions for processing transactions of the target objects are determined through the random search variable selection strategy, which avoids the low efficiency of determining the influence factors of the financial institutions for processing transactions of the target objects according to all possible combinations of variables in the related art, thereby achieving the technical effect of improving the efficiency of determining the influence factors of the financial institutions for processing transactions of the target objects.
[0192] Those skilled in the art can understand that Figure 5 The structure shown is only schematic, and the electronic device can also be a smart phone, a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, or other terminal devices. Figure 5 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 5 For example, the electronic device can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 5 For example, the electronic device can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure.
[0193] Those skilled in the art can understand that all or part of the steps of various methods in the above embodiments can be completed by instructing the terminal device related hardware through a program, and the program can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0194] Embodiment Four
[0195] The embodiments of the present application further provide a storage medium. Optionally, in the embodiments, the storage medium can be used to store the program code executed by the transaction data processing method provided in the first embodiment.
[0196] Optionally, in the embodiments, the storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0197] The present application further provides a computer program product, which, when executed on a data processing device, is adapted to execute the steps of the transaction data processing method.
[0198] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.
[0199] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0200] In the several embodiments provided by the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the unit embodiment described above is only schematic. For example, the division of the units is only a logical function division. There can be another division manner for actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, units or modules, and can be electrical or other forms.
[0201] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e. can be located in one place, or can be distributed on multiple network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the embodiments.
[0202] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0203] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in the form of a contribution to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes various media that can store program codes, such as a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, etc.
[0204] The above is only the preferred embodiment of the present application, and it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A method for processing transaction data, characterized in that: include: Obtain transaction data of N target objects to obtain a transaction data set, where N is a positive integer; Establishing a target regression model based on the transaction data set, wherein the type of the target regression model includes: a linear regression model; Based on the random search variable selection strategy and the target regression model, a target factor set is determined, wherein the target factor set includes: factors influencing the target object's selection of a financial institution to process a transaction.
2. The processing method according to claim 1, characterized in that The random search variable selection strategy includes: a variable selection strategy based on a Bayesian model, and determining a target factor set based on the random search variable selection strategy and the target regression model, including: Using a Bayesian model to fit the parameters of the target regression model to obtain a fitting result; Selecting variables of the target regression model based on the fitting results to obtain a selection result; The target factor set is determined based on the variables in the selection result.
3. The processing method according to claim 2, characterized in that The parameters of the target regression model include: the coefficients of the independent variables in the target regression model; the parameters of the target regression model are fitted using a Bayesian model to obtain a fitting result, including: Extracting the coefficients of the independent variables in the target regression model to obtain the parameters of the target regression model; Determining initial parameters of the Bayesian model based on the parameters of the target regression model, wherein the initial parameters include: parameters, indicator variables, and variance of the target regression model, and the indicator variables are used to select variables of the target regression model; The initial parameters of the Bayesian model are iteratively updated M times to obtain the fitting result, wherein, M is a positive integer.
4. The processing method according to claim 3, characterized in that The initial parameters of the Bayesian model are iteratively updated M times to obtain the fitting result, including: Determining N prior distributions of the Bayesian model, wherein the N prior distributions include: a prior distribution of parameters of the target regression model, a prior distribution of the indicative variable, and a prior distribution of variance; Performing M iterative updates on the initial parameters of the Bayesian model based on the N prior distributions of the Bayesian model, and determining the posterior distribution of the indicative variable; The fitting result is determined based on the posterior distribution of the indicative variable.
5. The processing method according to claim 4, characterized in that: The distribution type of the prior distribution of the parameters of the target regression model includes: mixed normal distribution, the distribution type of the prior distribution of the indicative variable includes: binomial distribution, and the distribution type of the prior distribution of the variance includes: inverse gamma distribution.
6. The processing method according to claim 3, characterized in that: The indicative variables include: S variables, each variable is associated with the coefficient of one of the independent variables in the target regression model, S is a positive integer, the fitting result includes: the iteratively updated indicative variables, and the variables of the target regression model are selected based on the fitting result to obtain a selection result, including: When a variable in the iteratively updated indicative variables is a first preset value, eliminating the independent variable associated with the variable in the target regression model to obtain a first processing result; When a variable in the iteratively updated indicative variables is a second preset value, retaining the independent variable associated with the variable in the target regression model to obtain a second processing result; The selection result is determined based on the first processing result and the second processing result.
7. The processing method according to claim 1, characterized in that Establishing a target regression model based on the transaction dataset includes: Establishing an initial regression model, wherein the initial regression model includes: S independent variables and a coefficient to be solved for each independent variable; Solving the coefficient of each independent variable in the initial regression model based on the transaction data to obtain a solution result; Based on the solution result, the target regression model is determined.
8. A transaction data processing device, characterized in that: include: an acquisition unit, configured to acquire transaction data of N target objects to obtain a transaction data set, where N is a positive integer; An establishing unit, configured to establish a target regression model based on the transaction data set, wherein the target regression model may be a linear regression model; The determination unit is configured to determine a target factor set based on a random search variable selection strategy and the target regression model, wherein the target factor set includes factors influencing the target object's selection of a financial institution for processing transactions.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the transaction data processing method according to any one of claims 1 to 7.
10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the transaction data processing method according to any one of claims 1 to 7 are implemented.