Univariate processing method and variable screening method
By exchanging encrypted parameters to construct and analyze one-dimensional linear regression models, the method addresses the lack of linear correlation assessment in federated learning, enhancing feature selection accuracy and security.
Patent Information
- Application Number
- CN202210418824.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-04-20
AI Technical Summary
In the federated learning scenario, the existing technology lacks a description of the degree of linear correlation between independent variables and dependent variables, resulting in the linear correlation between independent variables and dependent variables being unable to be effectively expressed, affecting the accuracy and interpretability of the model.
By passing encrypted intermediate parameters between different data owners, the linear correlation between independent variables and dependent variables is calculated, the independent variable's interpretation ability of dependent variables is analyzed using a one-variable linear regression model and the variance expansion coefficient R2, and the stable model is processed through binning to ensure data security.
On the premise of ensuring data security, the effective analysis of the linear correlation between independent variables and dependent variables provides a basis for the screening of subsequent characteristic variables and improves the explanatory and accuracy of the model.
Smart Images

Figure CN114692089B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a single variable processing method and a variable screening method. Background Art
[0002] The federated learning framework is a distributed artificial intelligence model training framework, which can help different data owners to achieve federated modeling and federated training without sharing private data, and can effectively solve the problems of data security and data islands.
[0003] Feature engineering is the most important part in machine learning modeling, which refers to the process of processing raw data into model training data. It generally includes three steps: feature preprocessing, feature selection, and feature dimensionality reduction. Among them, in feature selection, the feature univariate analysis method is adopted to analyze the distribution of each feature and its predictive ability for the label. The univariate analysis in the federated scenario includes the Weight of Evidence (WOE) and Information Value (IV).
[0004] However, indicators such as WOE and IV cannot represent the linear correlation degree between independent variables and dependent variables in the federated scenario. Summary of the Invention
[0005] The present invention provides a single variable processing method and a variable screening method to solve the lack of description of the linear correlation degree between independent variables and dependent variables in the existing technology.
[0006] In a first aspect, the present invention provides a univariate processing method, which is applied to a first data terminal in a univariate processing system. The univariate processing system includes the first data terminal and a second data terminal. Among them, the first data terminal stores a dependent variable, and the second data terminal stores an independent variable. The method includes: obtaining the difference between the dependent variable and the mean value of the dependent variable, and sending the difference to the second data terminal. The difference is used to calculate the regression coefficient of a univariate linear regression model constructed by the independent variable and the dependent variable; receiving a third parameter and an encrypted fourth parameter sent by the second data terminal, where the third parameter is calculated based on the regression coefficient and the mean value of the independent variable, and the fourth parameter is calculated based on the regression coefficient and the independent variable; obtaining the constant term of the univariate linear regression model according to the third parameter and the mean value of the dependent variable; obtaining an encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term, and the dependent variable, and obtaining a sum of squared deviations according to the difference between the dependent variable and the mean value of the dependent variable; obtaining an encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and sending the encrypted first correlation coefficient to the second data terminal for decryption; receiving the decrypted first correlation coefficient sent by the second data terminal, and outputting the first correlation coefficient.
[0007] As an optional embodiment, the step of obtaining the difference between the dependent variable and the mean value of the dependent variable, and sending the difference to the second data terminal, where the difference is used to calculate the regression coefficient of a univariate linear regression model constructed by the independent variable and the dependent variable, includes: receiving an encrypted first parameter sent by the second data terminal, where the first parameter is calculated based on the independent variable and the mean value of the independent variable; obtaining an encrypted second parameter according to the encrypted first parameter and the difference between the dependent variable and the mean value of the dependent variable, and sending the encrypted second parameter to the second data terminal for decryption. The decrypted second parameter is used to calculate the regression coefficient of a univariate linear regression model constructed by the independent variable and the dependent variable.
[0008] As an optional embodiment, the second data terminal includes a second key pair, where the second key pair includes a second public key and a second private key; among them, the encrypted fourth parameter and the encrypted first parameter are both encrypted by the second public key; the decrypted first correlation coefficient and the decrypted second parameter are both decrypted by the second private key.
[0009] As an alternative embodiment, the independent variables stored at the second data terminal are subjected to binning processing; before obtaining the difference between the dependent variable and the mean value of the dependent variable, the method further includes: sending the sample identifier and the encrypted sample tag value corresponding to the sample identifier to the second data terminal; receiving the encrypted sample tag statistical value sent by the second data terminal, where the encrypted sample tag statistical value is obtained by the second data terminal through statistical analysis of the encrypted sample tag values of each bin according to the sample identifier; decrypting the encrypted sample tag statistical value to obtain the dependent variable.
[0010] As an alternative embodiment, before sending the sample identifier and the encrypted sample tag value corresponding to the sample identifier to the second data terminal, the method further includes: generating a first key pair, where the first key pair includes a first public key and a first private key; encrypting the sample tag value with the first public key to obtain the encrypted sample tag value; the decrypting the encrypted sample tag statistical value includes: decrypting the encrypted sample tag statistical value with the first private key.
[0011] As an alternative embodiment, if the second data terminal includes multiple independent variables, the step of obtaining the difference between the dependent variable and the mean value of the dependent variable is iteratively executed until the first correlation coefficient corresponding to each independent variable is output.
[0012] As an alternative embodiment, after outputting the first correlation coefficient corresponding to each independent variable, the method further includes: selecting the independent variables whose first correlation coefficients meet a first preset condition to form a candidate data set.
[0013] In a second aspect, the present invention provides another single-variable processing method, which is applied to the second data terminal in a single-variable processing system. The single-variable processing system includes a first data terminal and the second data terminal. The first data terminal stores a dependent variable, and the second data terminal stores independent variables. The method includes: receiving the difference between the dependent variable and the mean value of the dependent variable sent by the first data terminal, and calculating the regression coefficient of the unary linear regression model constructed by the independent variable and the dependent variable according to the difference; obtaining a third parameter according to the regression coefficient and the mean value of the independent variable, obtaining a fourth parameter according to the regression coefficient and the independent variable, and encrypting the fourth parameter; sending the third parameter and the encrypted fourth parameter to the first data terminal, where the third parameter and the encrypted fourth parameter are used to calculate the encrypted first correlation coefficient corresponding to the independent variable; receiving the encrypted first correlation coefficient sent by the first data terminal, decrypting the encrypted first correlation coefficient, sending the decrypted first correlation coefficient to the first data terminal, and outputting the first correlation coefficient.
[0014] As an optional embodiment, receiving the difference between the dependent variable sent by the first data terminal and the mean value of the dependent variable, and calculating the regression coefficient of the unary linear regression model constructed by the independent variable and the dependent variable according to the difference includes: calculating a first parameter based on the independent variable and the mean value of the independent variable, encrypting the first parameter, and sending the encrypted first parameter to the first data terminal; receiving the encrypted second parameter sent by the first data terminal, where the encrypted second parameter is calculated based on the encrypted first parameter and the difference between the dependent variable and the mean value of the dependent variable; decrypting the encrypted second parameter, and calculating the regression coefficient of the unary linear regression model constructed by the independent variable and the dependent variable according to the decrypted second parameter.
[0015] Before receiving the difference between the dependent variable sent by the first data terminal and the mean value of the dependent variable, as an optional embodiment, it further includes: performing binning processing on the independent variable; receiving the sample identifier sent by the first data terminal and the encrypted sample tag value corresponding to the sample identifier; statistically analyzing the encrypted sample tag values of each bin according to the sample identifier to obtain the encrypted sample statistical value, and sending the encrypted sample statistical value to the first data terminal for decryption to obtain the dependent variable.
[0016] In a third aspect, the present invention provides a first data terminal, including a first processing module, a first sending module and a first receiving module: wherein, the first processing module is configured to obtain the difference between the dependent variable and the mean value of the dependent variable, and send the difference to the second data terminal through the first sending module, and the difference is used to calculate the regression coefficient of the unary linear regression model constructed by the independent variable and the dependent variable; the first receiving module is configured to receive the third parameter and the encrypted fourth parameter sent by the second data terminal, where the third parameter is calculated based on the regression coefficient and the mean value of the independent variable, and the fourth parameter is calculated based on the regression coefficient and the independent variable; the first processing module is further configured to obtain the constant term of the unary linear regression model according to the third parameter and the mean value of the dependent variable; obtain the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term and the dependent variable, and obtain the sum of squared deviations according to the difference between the dependent variable and the mean value of the dependent variable; obtain the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and send the encrypted first correlation coefficient to the second data terminal through the first sending module for decryption; the first receiving module is further configured to receive the decrypted correlation coefficient sent by the second data terminal and output the correlation coefficient.
[0017] Fourth aspect, the present invention provides a second data terminal, including a second processing module, a second sending module and a second receiving module; wherein, the second receiving module is configured to receive the difference between the dependent variable and the mean value of the dependent variable sent by the first data terminal; the second processing module is configured to calculate the regression coefficient of the univariate linear regression model constructed by the independent variable and the dependent variable according to the difference; obtain a third parameter according to the regression coefficient and the mean value of the independent variable, obtain a fourth parameter according to the regression coefficient and the independent variable, and perform encryption processing on the fourth parameter; the second sending module is configured to send the third parameter and the encrypted fourth parameter to the first data terminal, and the third parameter and the encrypted fourth parameter are used to calculate the encrypted correlation coefficient corresponding to the independent variable; the second receiving module is further configured to receive the encrypted first correlation coefficient sent by the first data terminal, decrypt the encrypted first correlation coefficient through the second processing module, and send the decrypted first correlation coefficient to the first data terminal through the second sending module.
[0018] Fifth aspect, the present invention provides a univariate processing system, including a first data terminal and a second data terminal; wherein, the first data terminal is configured to execute the method according to any one of the first aspect, and the second data terminal is configured to execute the method according to any one of the second aspect.
[0019] Sixth aspect, the present invention provides an electronic device, characterized in that it includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus; the memory is used to store a computer program; the processor is configured to, when executing the program stored on the memory, implement the method according to any one of the first aspect.
[0020] Seventh aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements the method according to any one of the first aspect.
[0021] In an eighth aspect, the present invention provides a variable screening method, which is applied to a variable screening system. The variable screening system includes a fifth data terminal and a sixth data terminal. Among them, the fifth data terminal stores a dependent variable and a first independent variable set, and the sixth data terminal stores a second independent variable set. The method includes: the fifth data terminal obtains a first correlation coefficient between each independent variable in the first independent variable set and the dependent variable; the fifth data terminal obtains the difference between the dependent variable and the mean value of the dependent variable, and sends the difference to the sixth data terminal; the sixth data terminal receives the difference, and calculates the regression coefficient of a univariate linear regression model constructed by any independent variable in the first independent variable set and the dependent variable according to the difference; the sixth data terminal obtains a third parameter according to the regression coefficient and the mean value of the independent variable, obtains a fourth parameter according to the regression coefficient and the independent variable, encrypts the fourth parameter, and sends the third parameter and the encrypted fourth parameter to the first data terminal; the fifth data terminal receives the third parameter and the encrypted fourth parameter, obtains the constant term of the univariate linear regression model according to the third parameter and the mean value of the dependent variable, obtains the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term and the dependent variable, and obtains the sum of squared deviations according to the difference between the dependent variable and the mean value of the dependent variable; obtains the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and sends the encrypted first correlation coefficient to the sixth data terminal; the sixth data terminal receives the encrypted first correlation coefficient, decrypts the encrypted first correlation coefficient, and sends the decrypted first correlation coefficient to the fifth data terminal; the fifth data terminal receives the decrypted first correlation coefficient and outputs the first correlation coefficient; iteratively execute the step of obtaining the difference between the dependent variable and the mean value of the dependent variable until the first correlation coefficient corresponding to each independent variable in the second independent variable set is output; select the independent variables in the first independent variable set whose first correlation coefficients meet the first preset condition to form a first candidate data set, and select the independent variables in the second independent variable set whose first correlation coefficients meet the first preset condition to form a second candidate data set; select any independent variable in the first candidate data set as the target variable, the other independent variables in the first candidate data set as the first input variables, and the second candidate data set as the second input variables; the sixth data terminal obtains the calculated value of the second input variables according to the second input variables and the second model parameters, and sends the calculated value of the second input variables to the fifth data terminal; the fifth data terminal obtains the predicted value of the target variable according to the calculated value of the second input variables, the first input variables and the first model parameters; obtains the sum of squared residuals according to the target variable and the predicted value of the target variable, and obtains the sum of squared deviations according to the target variable and the mean value of the target variable; determines the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and outputs the second correlation coefficient;Iteratively execute the step of selecting any independent variable in the first candidate dataset as the target variable until the second correlation coefficient corresponding to each independent variable in the first candidate dataset when it is used as the target variable is output; select any independent variable in the second candidate dataset as the target variable, the other independent variables in the second candidate dataset as the third input variables, and the first candidate dataset as the fourth input variables; the fifth data terminal obtains the calculated value of the fourth input variables according to the fourth input variables and the fourth model parameters, and sends the calculated value of the fourth input variables to the sixth data terminal; the sixth data terminal obtains the predicted value of the target variable according to the calculated value of the fourth input variables, the third input variables and the third model parameters; obtain the sum of squared residuals according to the target variable and the predicted value of the target variable, and obtain the total sum of squares according to the target variable and the mean value of the target variable; determine the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the total sum of squares, and output the second correlation coefficient; iteratively execute the step of selecting any independent variable in the second candidate dataset as the target variable until the second correlation coefficient corresponding to each independent variable in the second candidate dataset when it is used as the target variable is output; select the independent variables in the first candidate dataset whose second correlation coefficients meet the second preset condition to form a third candidate dataset, select the independent variables in the second candidate dataset whose second correlation coefficients meet the second preset condition to form a fourth candidate dataset, and the third candidate dataset and the fourth candidate dataset form the final candidate dataset.
[0022] In a ninth aspect, the present invention provides another variable screening method, which is applied to a variable screening system. The variable screening system includes a fifth data terminal and a sixth data terminal. Among them, the fifth data terminal stores a dependent variable and a first independent variable set, and the sixth data terminal stores a second independent variable set; the method includes: selecting any independent variable in the first independent variable set as a target variable, other independent variables in the first independent variable set as first input variables, and the second independent variable set as second input variables; the sixth data terminal obtains calculated values of the second input variables according to the second input variables and second model parameters, and sends the calculated values of the second input variables to the fifth data terminal; the fifth data terminal obtains predicted values of the target variable according to the calculated values of the second input variables, the first input variables, and first model parameters; obtaining the sum of squared residuals according to the target variable and the predicted value of the target variable, and obtaining the sum of squared deviations according to the target variable and the mean value of the target variable; determining the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and outputting the second correlation coefficient; iteratively executing the step of selecting any independent variable in the first independent variable set as the target variable until the second correlation coefficients corresponding to each independent variable in the first independent variable set as the target variable are output; selecting any independent variable in the second independent variable set as the target variable, other independent variables in the second independent variable set as third input variables, and the second independent variable set as fourth input variables; the fifth data terminal obtains calculated values of the fourth input variables according to the fourth input variables and fourth model parameters, and sends the calculated values of the fourth input variables to the sixth data terminal; the sixth data terminal obtains predicted values of the target variable according to the calculated values of the fourth input variables, the third input variables, and third model parameters; obtaining the sum of squared residuals according to the target variable and the predicted value of the target variable, and obtaining the sum of squared deviations according to the target variable and the mean value of the target variable; determining the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and outputting the second correlation coefficient; iteratively executing the step of selecting any independent variable in the second independent variable set as the target variable until the second correlation coefficients corresponding to each independent variable in the second independent variable set as the target variable are output; selecting the independent variables in the first independent variable set whose second correlation coefficients meet the second preset condition to form a fifth candidate data set, and selecting the independent variables in the second independent variable set whose second correlation coefficients meet the second preset condition to form a sixth candidate data set; the fifth data terminal obtains the first correlation coefficient between each independent variable in the fifth candidate data set and the dependent variable; the fifth data terminal obtains the difference between the dependent variable and the mean value of the dependent variable, and sends the difference to the sixth data terminal; the sixth data terminal receives the difference and calculates the regression coefficient of the simple linear regression model constructed by any independent variable in the sixth candidate data set and the dependent variable according to the difference.The sixth data terminal obtains a third parameter according to the regression coefficient and the independent variable mean, obtains a fourth parameter according to the regression coefficient and the independent variable, performs encryption processing on the fourth parameter, and sends the third parameter and the encrypted fourth parameter to the fifth data terminal; the fifth data terminal receives the third parameter and the encrypted fourth parameter, obtains the constant term of the unary linear regression model according to the third parameter and the dependent variable mean, obtains the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term and the dependent variable, and obtains the sum of squared deviations according to the difference between the dependent variable and the dependent variable mean; obtains the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and sends the encrypted first correlation coefficient to the sixth data terminal; the sixth data terminal receives the encrypted first correlation coefficient, decrypts the encrypted first correlation coefficient, and sends the decrypted first correlation coefficient to the fifth data terminal; the fifth data terminal receives the decrypted first correlation coefficient and outputs the first correlation coefficient; iteratively execute the step of obtaining the difference between the dependent variable and the dependent variable mean until the first correlation coefficients corresponding to each independent variable in the sixth candidate data set are output; select the independent variables whose first correlation coefficients in the fifth candidate data set satisfy the first preset condition to form a seventh candidate data set, and select the independent variables whose first correlation coefficients in the sixth candidate data set satisfy the first preset condition to form an eighth candidate data set, and the seventh candidate data set and the eighth candidate data set form the final candidate data set.
[0023] In a tenth aspect, the present invention provides a variable screening system, including a fifth data terminal and a sixth data terminal; wherein, the fifth data terminal and the sixth data terminal are used to execute the method described in the eighth aspect or the ninth aspect.
[0024] At least some of the above technical solutions provided by the embodiments of the present invention have the following advantages:
[0025] By transmitting intermediate parameters between different data owners, it is possible to effectively analyze the linear correlation degree between independent variables and dependent variables while ensuring the data security of both parties, providing an effective basis for the screening of candidate feature variables. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 A schematic diagram of a system architecture provided by an embodiment of the present invention;
[0029] Figure 2 A schematic flow diagram of a single variable processing method provided by Embodiment 1 of the present invention;
[0030] Figure 3 A detailed implementation flow diagram of steps S101 and S102 in Embodiment 1 of the present invention;
[0031] Figure 4 A schematic flow diagram of a single variable processing method provided by Embodiment 2 of the present invention;
[0032] Figure 5 A schematic diagram of a unary linear regression model fitted by a single variable and a dependent variable provided by an embodiment of the present invention;
[0033] Figure 6 A schematic flow diagram of a single variable processing method provided by Embodiment 3 of the present invention;
[0034] Figure 7 A schematic diagram of the structure of a first data terminal provided by an embodiment of the present invention;
[0035] Figure 8 A schematic diagram of the structure of a second data terminal provided by an embodiment of the present invention;
[0036] Figure 9 A schematic flow diagram of a multi-variable processing method provided by Embodiment 4 of the present invention;
[0037] Figure 10 A schematic diagram of the variance inflation factor VIF corresponding to each feature variable provided by an embodiment of the present invention;
[0038] Figure 11 A schematic flow diagram of a multi-variable processing method provided by Embodiment 5 of the present invention;
[0039] Figure 12 A schematic flow diagram of a multi-variable processing method provided by Embodiment 6 of the present invention;
[0040] Figure 13 A schematic diagram of the structure of a third data terminal provided by an embodiment of the present invention;
[0041] Figure 14 A schematic diagram of the structure of a fourth data terminal provided by an embodiment of the present invention;
[0042] Figure 15 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation mode
[0043] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] The federated learning framework is a distributed artificial intelligence model training framework, which can help different data owners achieve federated modeling and federated training without sharing private data, and can effectively solve the problems of data security and data islands.
[0045] Feature engineering is the most important part in machine learning modeling, which refers to the process of processing raw data into model training data, aiming to extract the information of the data to the greatest extent for the model to use. It generally includes the following steps:
[0046] (1) Feature preprocessing: The raw data is probably "dirty and messy", and abnormal feature sample processing, missing value processing, standardization and normalization, etc. need to be carried out;
[0047] (2) Feature selection: After the feature preprocessing is completed, meaningful features need to be screened out and input into the model for training. The screening methods include univariate feature analysis, that is, analyzing the distribution of each feature and its prediction ability for the label. Common indicators include WOE, IV, PSI, KS, etc. This method is widely used in the scoring card modeling in the consumer finance field.
[0048] (3) Feature dimensionality reduction: After the feature selection is completed, due to the large feature matrix, problems such as low calculation efficiency and high model complexity may occur, which can be solved by feature dimensionality reduction. Specific methods include PCA (Principal Component Analysis), LDA (Linear Discriminant Analysis), ICA (Independent Component Analysis), etc.
[0049] The univariate analysis in the federated scenario includes Weight of Evidence (WOE for short) and Information Value (IV for short). However, both WOE and IV are used to select relatively important variables to be added to the model, and the prediction strength can be used as the basis for judging whether a variable is important. The linear correlation between independent variables and dependent variables is lacking, and there is also a lack of a measure of the severity of multicollinearity in the multiple linear regression model (multiple). If there is multicollinearity among features, the weight parameter estimation of the model will be distorted or difficult to estimate accurately.
[0050] In view of the above technical problems, the technical concept of the present invention lies in that different data owners calculate the variance inflation factor R by transmitting encrypted intermediate parameters to each other. 2 The variance inflation factor VIF is calculated to analyze the explanatory power of the explanatory variable for the dependent variable, and the variance inflation factor VIF calculated is used to analyze the multicollinearity among features, thereby increasing the interpretability of the variables included in the model.
[0051] Figure 1 FIG. shows a schematic diagram of a system architecture provided by an embodiment of the present invention. Figure 1 As shown, the system architecture 100 includes: an A-side server 101, a B-side server 102, and a network 103. Among them, the A-side server 101 and the B-side server 102 each store part of the sample data, and encrypted intermediate parameter transmission is realized through the network 103. The network 103 can be a wired or wireless communication link, an optical fiber cable, or the like.
[0052] It should be noted that the A-side server 101 and the B-side server 102 cooperate together to implement the following various embodiments. The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0053] The present invention provides a univariate processing method, which is applied to a univariate processing system. The univariate processing system includes a first data end and a second data end. Among them, the first data end stores a dependent variable, and the second data end stores an independent variable. Reference can be made to Figure 1 shown (as Figure 1 in which the A-side server is equivalent to the first data end, and the B-side server is equivalent to the second data end).
[0054] Figure 2 FIG. shows a schematic flowchart of a univariate processing method provided by Embodiment 1 of the present invention. Figure 2 As shown, the univariate processing method includes:
[0055] Step S101: The first data end obtains the difference between the dependent variable and the mean value of the dependent variable, and sends the difference to the second data end.
[0056] Among them, the difference is used to calculate the regression coefficient of the univariate linear regression model constructed by the independent variable and the dependent variable.
[0057] Combined with Figure 1 it can be said that the A-side server stores the dependent variable y corresponding to n sample sizes, and the B-side server stores the independent variable x corresponding to n sample sizes. t , Optionally, the A-side or the B-side also stores other independent variables. When it is necessary to analyze the explanatory power of the univariate x t for the dependent variable y, regression calculation is realized by the least squares method, that is, the univariate linear regression model constructed by the independent variable and the dependent variable is as shown in formula (1):
[0058] y' = f(x t,i ) = a0 + a1x t (1)
[0059] Among them, y' represents the predicted value of the dependent variable, and x t,i represents the i-th sample value of the independent variable x t ; a1 represents the regression coefficient of the simple linear regression model, and its calculation formula is as shown in (2):
[0060]
[0061] Among them, represents the mean value of x t ; y i represents the i-th sample label value, represents the mean value of y.
[0062] a0 represents the constant term of the simple linear regression model, and its calculation formula is as shown in (3):
[0063]
[0064] Next, the linear correlation degree of the independent variable to the dependent variable can be calculated according to the predicted value of the dependent variable, that is, the variance inflation factor R 2 , as shown in formula (4):
[0065]
[0066] Among them, RMSE represents the sum of squared residuals of the constructed regression model, and its calculation formula is as shown in (5); SST represents the sum of squared deviations of the constructed regression model, and its calculation formula is as shown in (6).
[0067]
[0068]
[0069] In this step, end A can obtain the mean value of the dependent variable according to y i and then calculate to obtain the difference between y and i and and send to end B. It should be noted that what end A sends to end B is the difference , and end B cannot reverse calculate y from the difference, ensuring the data security of end A. i
[0070] Step S102: The second data terminal receives the difference between the dependent variable and the mean value of the dependent variable sent by the first data terminal, and calculates the regression coefficient of the unary linear regression model constructed by the independent variable and the dependent variable according to the difference.
[0071] Specifically, after the B terminal receives it, the regression coefficient a1 of the unary linear regression model can be calculated according to formula (2).
[0072] Figure 3 FIG. is a detailed implementation process schematic diagram of steps S101 and S102 in Embodiment 1 of the present invention. As Figure 3 shown, steps S101 and S102 include steps S1011-S1013, as follows:
[0073] Step S1011: The second data terminal calculates a first parameter based on the independent variable and the mean value of the independent variable, encrypts the first parameter, and sends the encrypted first parameter to the first data terminal.
[0074] Correspondingly, on the side of the first data terminal, the first data terminal receives the encrypted first parameter sent by the second data terminal, and the first parameter is calculated based on the independent variable and the mean value of the independent variable.
[0075] In this step, the B terminal can calculate the mean value of the independent variable according to x t,i and then calculate the first parameter according to x t,i The first parameter u1 is related to obtaining the regression coefficient. The B terminal encrypts the first parameter u1 and then sends the encrypted u1 to the first data terminal.
[0076] Optionally, the second data terminal includes a second key pair, and the second key pair includes a second public key and a second private key; the encrypted first parameter is obtained by encrypting with the second public key. Specifically, the B terminal server generates a second key pair, including a second public key PK B and a second private key SK B . After the B terminal obtains u1, it encrypts u1 with PK B to obtain the encrypted first parameter Enc B (u1), and sends Enc B (u1) to the first data terminal. It should be noted that Enc B (u1) is encrypted with PK B and can only be decrypted with the second private key SK of the B terminal B . That is to say, the A terminal cannot reverse-decrypt u1.
[0077] Step S1012: The first data terminal obtains an encrypted second parameter based on the encrypted first parameter and the difference between the dependent variable and the mean of the dependent variables, and sends the encrypted second parameter to the second data terminal.
[0078] Correspondingly, on the side of the second data terminal, it receives the encrypted second parameter sent by the first data terminal, where the encrypted second parameter is calculated based on the encrypted first parameter and the difference between the dependent variable and the mean of the dependent variables.
[0079] In this step, end A obtains the encrypted second parameter through calculation It should be noted that since Enc B (u1) is encrypted by PK B , therefore, the obtained second parameter v1 is also encrypted by PK B .
[0080] Step S1013: The second data terminal decrypts the encrypted second parameter and calculates the regression coefficient of the unary linear regression model constructed by the independent variable and the dependent variable based on the decrypted second parameter.
[0081] Optionally, the decrypted second parameters are all obtained by decrypting with the second private key, that is, end B uses the second private key SK B to decrypt Enc B (v1) to obtain the decrypted second parameter v1, and then end B calculates the regression coefficient a1 according to v1. The calculation formula is as shown in (7), and formula (7) is a deformation of formula (2).
[0082]
[0083] Step S103: The second data terminal obtains a third parameter based on the regression coefficient and the mean of the independent variables, obtains a fourth parameter based on the regression coefficient and the independent variable, and performs an encryption process on the fourth parameter, and sends the third parameter and the encrypted fourth parameter to the first data terminal.
[0084] Among them, the third parameter and the encrypted fourth parameter are used to calculate the encrypted first correlation coefficient corresponding to the independent variable. Correspondingly, on the side of the first data terminal, it receives the third parameter and the encrypted fourth parameter sent by the second data terminal, where the third parameter is calculated based on the regression coefficient and the mean of the independent variables, and the fourth parameter is calculated based on the regression coefficient and the independent variable.
[0085] In this step, end B obtains the third parameter based on the regression coefficient a1 and the mean of the independent variables and obtains the fourth parameter based on the regression coefficient a1 and the independent variable x t to obtain the fourth parameter w1 = a1xt and encrypt w1, and send the encrypted w1 to end A. Optionally, the second data end uses the second public key PK B to encrypt w1 to obtain the encrypted fourth parameter Enc B (w1).
[0086] Optionally, the fifth parameter can also be obtained according to the regression coefficient a1 and the independent variable x t to obtain the fifth parameter and use the second public key PK B to encrypt the fifth parameter w2 to obtain Enc B (w2), and send Enc B (w2) to the first data end.
[0087] Step S104: The first data end obtains the constant term of the unary linear regression model according to the third parameter and the mean value of the dependent variable.
[0088] In this step, end A obtains the constant term of the unary linear regression model according to u2 and which is a deformation of formula (3). It is a deformation of formula (3).
[0089] Step S105: The first data end obtains the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term, and the dependent variable, and obtains the sum of squared deviations according to the difference between the dependent variable and the mean value of the dependent variable.
[0090] In this step, the first data end obtains the encrypted root mean square error of the sum of squared residuals RMSE according to Enc B (w1), a0, and y i , and optionally, obtains the encrypted RMSE according to Enc B (w1), Enc B (w2), a0, and y i , and its calculation formula is as shown in (8), which is a deformation of formula (5); the first data end also obtains the sum of squared deviations SST according to y i , , and the calculation formula is as shown in (6).
[0091]
[0092] Step S106: The first data end obtains the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and sends the encrypted first correlation coefficient to the second data end.
[0093] In this step, the first correlation coefficient can also be called the variance inflation factor R 2 , and end A obtains it according to Enc B(RMSE), SST obtains the encrypted variance inflation factor Send Enc B (R 2 ) to end B.
[0094] Step S107: The second data terminal receives the encrypted first correlation coefficient sent by the first data terminal, decrypts the encrypted first correlation coefficient, sends the decrypted first correlation coefficient to the first data terminal, and outputs the first correlation coefficient.
[0095] In this step, end B uses the second private key SK B to encrypt Enc B (R 2 ) to obtain the decrypted first correlation coefficient, that is, the variance inflation factor R 2 , and sends R 2 to end A. At the same time, end B outputs R 2 .
[0096] Step S108: The first data terminal receives the decrypted first correlation coefficient sent by the second data terminal and outputs the first correlation coefficient.
[0097] As an optional embodiment, if the second data terminal includes multiple independent variables, the step of obtaining the difference between the dependent variable and the mean of the dependent variable is iteratively executed until the first correlation coefficient corresponding to each independent variable is output.
[0098] Specifically, if end B includes independent variables in multiple dimensions, this embodiment can be used to analyze the first correlation coefficient between each independent variable and the dependent variable in the first data terminal.
[0099] As an optional embodiment, after outputting the first correlation coefficient corresponding to each independent variable, it further includes: selecting independent variables whose first correlation coefficient satisfies a first preset condition to form a candidate data set.
[0100] Specifically, the first correlation coefficient, that is, the variance inflation factor R 2 describes the magnitude of the effect of this independent variable on the dependent variable. If R 2 is small, it means that the dependent variable is less explained by this independent variable and the linear correlation degree is low. The first preset condition can be set to be greater than a certain threshold. For example, the threshold is 0.3, that is, only when R 2 is greater than 0.3, the degree of explanation of this independent variable to the dependent variable is relatively high, and this independent variable can be selected as the candidate data set. The candidate data set can be used as a training data set to train the model or used in other aspects.
[0101] The univariate processing method provided by the embodiment of the present invention obtains the difference between the dependent variable and the mean value of the dependent variable, and sends the difference to the second data terminal. The difference is used to calculate the regression coefficient of the unary linear regression model constructed by the independent variable and the dependent variable; receive the third parameter and the encrypted fourth parameter sent by the second data terminal, wherein the third parameter is calculated according to the regression coefficient and the mean value of the independent variable, and the fourth parameter is calculated according to the regression coefficient and the independent variable; obtain the constant term of the unary linear regression model according to the third parameter and the mean value of the dependent variable; obtain the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term and the dependent variable, and obtain the sum of squared deviations according to the difference between the dependent variable and the mean value of the dependent variable; obtain the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and send the encrypted first correlation coefficient to the second data terminal for decryption; receive the decrypted first correlation coefficient sent by the second data terminal, and output the first correlation coefficient; that is, in the embodiment of the present invention, by transmitting encrypted intermediate parameters between two data providers, while ensuring the data security of both parties, the linear correlation degree of a single variable to the dependent variable can be analyzed, providing an effective basis for subsequent screening of characteristic variables and the like.
[0102] Based on the above embodiment, Figure 4 As shown in the flowchart of a univariate processing method provided by the second embodiment of the present invention, in this embodiment, the independent variable stored in the second data terminal is processed by binning. As Figure 4 shown, the univariate processing method includes:
[0103] Step S201, the first data terminal sends the sample identifier and the encrypted sample tag value corresponding to the sample identifier to the second data terminal.
[0104] Correspondingly, on the second data terminal side, receive the sample identifier and the encrypted sample tag value corresponding to the sample identifier sent by the first data terminal.
[0105] Step S202, the second data terminal statistically processes the encrypted sample tag values of each bin according to the sample identifier to obtain the encrypted sample tag statistical value, and sends the encrypted sample tag statistical value to the first data terminal.
[0106] Correspondingly, on the first data terminal side, receive the encrypted sample tag statistical value sent by the second data terminal, where the encrypted sample tag statistical value is obtained by the second data terminal statistically processing the encrypted sample tag values of each bin according to the sample identifier.
[0107] Step S203, the first data terminal decrypts the encrypted sample tag statistical value to obtain the dependent variable.
[0108] Step S204: The first data terminal obtains the difference between the dependent variable and the mean of the dependent variable, and sends the difference to the second data terminal.
[0109] Wherein, the difference is used to calculate the regression coefficient of the simple linear regression model constructed by the independent variable and the dependent variable.
[0110] Step S205: The second data terminal receives the difference between the dependent variable and the mean of the dependent variable sent by the first data terminal, and calculates the regression coefficient of the simple linear regression model constructed by the independent variable and the dependent variable according to the difference.
[0111] Step S206: The second data terminal obtains a third parameter according to the regression coefficient and the mean of the independent variable, obtains a fourth parameter according to the regression coefficient and the independent variable, encrypts the fourth parameter, and sends the third parameter and the encrypted fourth parameter to the first data terminal.
[0112] Step S207: The first data terminal obtains the constant term of the simple linear regression model according to the third parameter and the mean of the dependent variable.
[0113] Step S208: The first data terminal obtains the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term, and the dependent variable, and obtains the sum of squared deviations according to the difference between the dependent variable and the mean of the dependent variable.
[0114] Step S209: The first data terminal obtains the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and sends the encrypted first correlation coefficient to the second data terminal.
[0115] Step S210: The second data terminal receives the encrypted first correlation coefficient sent by the first data terminal, decrypts the encrypted first correlation coefficient, sends the decrypted first correlation coefficient to the first data terminal, and outputs the first correlation coefficient.
[0116] Step S211: The first data terminal receives the decrypted first correlation coefficient sent by the second data terminal, and outputs the first correlation coefficient.
[0117] The implementation manners of steps S204 - S211 in this embodiment are respectively similar to those of steps S101 - S108 in the above embodiment, and will not be elaborated here.
[0118] The difference from the above embodiment is that, in order to reduce the risk of overfitting of the univariate linear regression model to be constructed and obtain a more stable regression model, in this embodiment, the independent variables stored in the second data terminal are subjected to binning processing; the sample identifier and the encrypted sample label value corresponding to the sample identifier are sent to the second data terminal; the encrypted sample label statistical value sent by the second data terminal is received, where the encrypted sample label statistical value is obtained by the second data terminal statistically processing the encrypted sample label values of each bin according to the sample identifier; the encrypted sample label statistical value is decrypted to obtain the dependent variable.
[0119] Specifically, first, all variables in the B end are binned, a total of m bins are divided, the i-th bin has a total of n i sample sizes, and the sample x t,i in the i-th bin and the dependent variable y i are taken for the mean value:
[0120] Then, the regression calculation is realized by using the least squares method, that is, the univariate linear regression model constructed by the independent variable and the dependent variable is as shown in formula (9):
[0121] Y′ = f(X t,i ) = a0 + a1X t (9)
[0122] where, X t = [X t,0 , X t,1 , ΛX t,m T , where,
[0123] Finally, calculate according to the regression prediction value where,
[0124] In this embodiment, the A end first sends the sample identifier and the encrypted sample label value y in the sample data to the second data terminal; then, the B end statistically processes the y of each bin according to the sample identifier to obtain the encrypted sample label statistical value y_bin_sum corresponding to each bin, and sends the encrypted y_bin_sum to the A end; the A end decrypts the encrypted y_bin_sum to obtain the decrypted y_bin_sum corresponding to each bin, and then uses to obtain the dependent variable Y i corresponding to each bin; then the A end and the B end cooperate to execute the method steps as described in Embodiment 1, so as to obtain the first correlation coefficient corresponding to the independent variable.
[0125] Optionally, before step 201, it further includes: the first data terminal generates a first key pair, the first key pair includes a first public key and a first private key; encrypts the sample tag value through the first public key to obtain the encrypted sample tag value; the decryption process of the encrypted sample tag statistical value in step 210 includes: decrypting the encrypted sample tag statistical value through the first private key.
[0126] Specifically, the A terminal generates a first key pair, including a first public key PK A and a first private key SK A , the A terminal uses PK A to encrypt the sample tag value y to obtain Enc A (y), and sends Enc A (y) and the sample identifier to the B terminal; then, the B terminal counts the encrypted sample tag values of each bin according to the sample identifier to obtain the encrypted sample tag statistical value Enc A (y_bin_sum), and sends Enc A (y_bin_sum) to the A terminal; the A terminal uses the second private key SK A to decrypt it to obtain the corresponding Y of each bin i .
[0127] Then, the A terminal obtains the dependent variable mean according to the dependent variable Y i Then calculates to obtain the dependent variable Y Then calculates the difference between the dependent variable Y i and the dependent variable mean and sends to the B terminal; so that the B terminal calculates the regression coefficient of the unary linear regression model constructed by the independent variable and the dependent variable according to ; then, the B terminal obtains the third parameter according to the regression coefficient a1 and the independent variable mean Obtains the fourth parameter w1 = a1X according to the regression coefficient a1 and the independent variable X t t , and encrypts w1, and sends the encrypted w1 to the A terminal. Optionally, B encrypts w1 through the second public key PK t B to obtain the encrypted fourth parameter Enc B B (w1). Optionally, the fifth parameter w2 = a1 can also be obtained according to the regression coefficient a1 and the independent variable X t 2 (X (X t X t ), and encrypts the fifth parameter w2 through the second public key PK B to obtain Enc B (w2), and sends EncB (w2) is sent to end A; then, end A, based on u2 and obtains the constant term of the univariate linear regression model Based on Enc B (w1), a0, and Y i obtains the encrypted root mean square error of residuals RMSE. Optionally, based on Enc B (w1), Enc B (w2), a0, and Y i obtains the encrypted RMSE, and its calculation formula is as shown in (10); the first data end also, based on Y i , obtains the total sum of squared deviations SST, and the calculation formula is as shown in (11).
[0128]
[0129]
[0130] End A, based on Enc B (RMSE), SST, obtains the encrypted variance inflation factor sends Enc B (R 2 ) to end B; end B, through the second private key SK B decrypts Enc B (R 2 ) to obtain the decrypted first correlation coefficient, that is, the variance inflation factor R 2 , and sends R 2 to end A, outputs R 2 , and at the same time, the second data end outputs R 2 .
[0131] Figure 5 is a schematic diagram of a univariate linear regression model fitted by a single variable and a dependent variable provided by an embodiment of the present invention. After fitting the univariate linear regression model as shown in Figure 5 , the variance inflation factor R 2 corresponding to the independent variable can be calculated.
[0132] The univariate processing method provided by the embodiment of the present invention performs binning processing on the independent variable stored at the second data end; sends the sample identifier and the encrypted sample tag value corresponding to the sample identifier to the second data end; receives the encrypted sample tag statistical value sent by the second data end, where the encrypted sample tag statistical value is obtained by the second data end through statistical analysis of the encrypted sample tag values of each bin according to the sample identifier; decrypts the encrypted sample tag statistical value to obtain the dependent variable; that is, by binning the independent variable, the linear regression model is made more stable, and by encrypting and sending the sample tag value to another data provider, the security of the data is ensured.
[0133] To further understand the embodiment of the present invention, Figure 6 FIG. is a schematic flow chart of a univariate processing method provided by Embodiment 3 of the present invention. In combination with Figure 1 and Figure 6 , the univariate processing method includes:
[0134] Step 1.A generates a pair of public and private keys PK A , SK A , and sends the public key PK A to B;
[0135] Step 2.B generates a pair of public and private keys PK B , SK B (PK A ≠PK B , SK A ≠SK B ), and sends the public key PK B to A;
[0136] Step 3.A sends the encrypted label Enc A (y) to B;
[0137] Step 4.B counts the number of positive and negative samples of each bin of its own features, and sends the label statistical information Enc A (y_bin_sum) to A;
[0138] Step 5.A uses the private key SK A to decrypt Enc A (y_bin_sum), and calculates the y i value at the B end;
[0139] Step 6.B calculates encrypts u1 with the public key PK B , and sends Enc B (u1) to A;
[0140] Step 7.A calculates Send Enc B (v1) to B;
[0141] Step 8. B decrypts Enc B (v1) using the private key SK B and calculates then sends u2 to A;
[0142] Step 9. B calculates w2 = a1X t , encrypts w1 and w2 using the public key PK B and sends Enc B (w1) and Enc B (w2) to A;
[0143] Step 10. A calculates
[0144]
[0145] then sends Enc B (R 2 ) to B;
[0146] Step 11. B decrypts Enc B (R B ) using the private key SK 2 and sends R 2 to A;
[0147] Step 12. Repeat Steps 6 to 11 until all independent variables at the B side have been analyzed.
[0148] Specifically, to ensure that the information transmitted between the two parties cannot be cracked by a third party, the server at the A side can encrypt the message sent to the server at the B side using the private key SK A , and the server at the B side can decrypt the message using the public key PK A received from the server at the A side; similarly, the server at the B side can encrypt the message sent to the server at the A side using SK B , and the server at the A side can decrypt the message using the public key PK B received from the server at the B side.
[0149] In summary, the server at the A side and the server at the B side each have different data. The two parties can perform data-related calculations locally, and only encrypted intermediate parameters are transmitted between the two parties. The receiving party cannot reverse-decrypt the original data, realizing the analysis of the linear correlation degree between each independent variable and dependent variable on the basis of ensuring the security of their respective data, providing an effective basis for subsequent variable screening, etc.
[0150] The embodiment of the present invention also provides a first data terminal.Figure 7 A structural schematic diagram of a first data terminal provided by an embodiment of the present invention is shown as Figure 7 shown. The first data terminal includes a first processing module 10, a first sending module 11, and a first receiving module 12;
[0151] Among them, the first processing module 10 is configured to obtain the difference between the dependent variable and the mean value of the dependent variable, and send the difference to the second data terminal through the first sending module 11. The difference is used to calculate the regression coefficient of a unary linear regression model constructed by the independent variable and the dependent variable. The first receiving module 12 is configured to receive a third parameter and an encrypted fourth parameter sent by the second data terminal. Among them, the third parameter is calculated based on the regression coefficient and the mean value of the independent variable, and the fourth parameter is calculated based on the regression coefficient and the independent variable. The first processing module 10 is further configured to obtain the constant term of the unary linear regression model based on the third parameter and the mean value of the dependent variable; obtain the encrypted sum of squared residuals based on the encrypted fourth parameter, the constant term, and the dependent variable, and obtain the sum of squared deviations based on the difference between the dependent variable and the mean value of the dependent variable; obtain the encrypted first correlation coefficient corresponding to the independent variable based on the encrypted sum of squared residuals and the sum of squared deviations, and send the encrypted first correlation coefficient to the second data terminal through the first sending module 11 for decryption; the first receiving module 12 is further configured to receive the decrypted first correlation coefficient sent by the second data terminal and output the first correlation coefficient.
[0152] As an optional embodiment, the first receiving module 12 is specifically configured to receive an encrypted first parameter sent by the second data terminal, and the first parameter is calculated based on the independent variable and the mean value of the independent variable. The first processing module 10 is specifically configured to obtain an encrypted second parameter based on the encrypted first parameter and the difference between the dependent variable and the mean value of the dependent variable, and send the encrypted second parameter to the second data terminal through the first sending module 11 for decryption. The decrypted second parameter is used to calculate the regression coefficient of a unary linear regression model constructed by the independent variable and the dependent variable.
[0153] As an optional embodiment, the second data terminal includes a second key pair, and the second key pair includes a second public key and a second private key. Among them, the encrypted fourth parameter and the encrypted first parameter are both encrypted by the second public key; the decrypted first correlation coefficient and the decrypted second parameter are both decrypted by the second private key.
[0154] As an alternative embodiment, the independent variables stored in the second data terminal are subjected to binning processing; the first sending module 11 is further configured to send a sample identifier and an encrypted sample tag value corresponding to the sample identifier to the second data terminal; the first receiving module 12 is configured to receive an encrypted sample tag statistical value sent by the second data terminal, where the encrypted sample tag statistical value is obtained by the second data terminal by statistically analyzing the encrypted sample tag values of each bin according to the sample identifier; the first processing module 10 is configured to perform decryption processing on the encrypted sample tag statistical value to obtain the dependent variable.
[0155] As an alternative embodiment, the first processing module 10 is further configured to generate a first key pair, where the first key pair includes a first public key and a first private key; encrypt the sample tag value through the first public key to obtain the encrypted sample tag value; perform decryption processing on the encrypted sample tag statistical value through the first private key.
[0156] As an alternative embodiment, if the second data terminal includes multiple independent variables, the step of obtaining the difference between the dependent variable and the mean value of the dependent variable is iteratively executed until the first correlation coefficient corresponding to each independent variable is output.
[0157] As an alternative embodiment, the first processing module 10 is further configured to select independent variables whose first correlation coefficient meets a first preset condition to form a candidate data set.
[0158] The first data terminal provided in this embodiment has a similar implementation principle and technical effect to the above embodiment, and will not be elaborated here.
[0159] An embodiment of the present invention further provides a second data terminal. Figure 8 Shown in Figure 8 is a schematic structural diagram of a second data terminal provided in an embodiment of the present invention. As shown, the second data terminal includes a second processing module 20, a second sending module 21, and a second receiving module 22;
[0160] Among them, the second receiving module 22 is configured to receive the difference between the dependent variable and the mean value of the dependent variable sent by the first data terminal; the second processing module 20 is configured to calculate the regression coefficient of the univariate linear regression model constructed by the independent variable and the dependent variable according to the difference; obtain a third parameter according to the regression coefficient and the mean value of the independent variable, obtain a fourth parameter according to the regression coefficient and the independent variable, and perform encryption processing on the fourth parameter; the second sending module 21 is configured to send the third parameter and the encrypted fourth parameter to the first data terminal, and the third parameter and the encrypted fourth parameter are used to calculate the encrypted correlation coefficient corresponding to the independent variable; the second receiving module 22 is further configured to receive the encrypted first correlation coefficient sent by the first data terminal, decrypt the encrypted first correlation coefficient through the second processing module 20, and send the decrypted first correlation coefficient to the first data terminal through the second sending module 21.
[0161] As an optional embodiment, the second processing module 20 is configured to calculate a first parameter according to the independent variable and the mean value of the independent variable, encrypt the first parameter, and send the encrypted first parameter to the first data terminal through the second sending module 21; the second receiving module 22 is configured to receive the encrypted second parameter sent by the first data terminal, where the encrypted second parameter is calculated according to the encrypted first parameter and the difference between the dependent variable and the mean value of the dependent variable; the second processing module 20 is further configured to decrypt the encrypted second parameter and calculate the regression coefficient of the univariate linear regression model constructed by the independent variable and the dependent variable according to the decrypted second parameter.
[0162] As an optional embodiment, the second processing module 20 is further configured to perform binning processing on the independent variable; the second receiving module 22 is further configured to receive the sample identifier sent by the first data terminal and the encrypted sample tag value corresponding to the sample identifier; the second processing module is further configured to count the encrypted sample tag values of each bin according to the sample identifier to obtain an encrypted sample tag statistical value, and send the encrypted sample tag statistical value to the first data terminal through the second sending module 21 for decryption to obtain the dependent variable.
[0163] The second data terminal provided in this embodiment has the same implementation principle and technical effects as those in the above embodiments, and will not be elaborated here.
[0164] The embodiment of the present invention further provides a univariate processing system. Refer to Figure 1 As shown, the univariate processing system includes a first data terminal and a second data terminal; among them, the first data terminal and the second data terminal are configured to execute the method described in any one of the above embodiments.
[0165] The univariate processing system provided in this embodiment has the same implementation principle and technical effects as those in the above embodiments, and will not be elaborated here.
[0166] The present invention also provides a multi-variable processing method, which is applied to a multi-variable processing system. The multi-variable processing system includes a third data terminal and a fourth data terminal. Among them, the first characteristic variable set is stored in the third data terminal, and the second characteristic variable set is stored in the second data terminal. For reference, Figure 1 as shown Figure 1 in which the A-side server is equivalent to the third data terminal, and the B-side server is equivalent to the fourth data terminal).
[0167] Figure 9 is a schematic flowchart of a multi-variable processing method provided in Embodiment 4 of the present invention. As Figure 9 shown, the multi-variable processing method includes:
[0168] Step S301: Select any characteristic variable in the first characteristic variable set as the target variable, other characteristic variables in the first characteristic variable set as the first input variables, and the second characteristic variable set as the second input variables.
[0169] Specifically, the first characteristic variable set is stored at the A side The has a feature dimension of f_dim A , select any one characteristic variable x from it as the target variable (dependent variable), that is t , and other characteristic variables in the first characteristic variable set are used as the first input variables, that is, the first input variables . The second characteristic variable set X is stored at the B side B , and the X B has a feature dimension of f_dim B , and the X B is used as the second input variable.
[0170] Step S302: The fourth data terminal obtains the calculated value of the second input variable according to the second input variable and the second model parameter, and sends the calculated value of the second input variable to the third data terminal.
[0171] Among them, the calculated value of the second input variable is used to determine the second correlation coefficient corresponding to the target variable. Correspondingly, on the side of the third data terminal, the calculated value of the second input variable sent by the fourth data terminal is received, where the calculated value of the second input variable is calculated according to the second input variable and the second model parameter.
[0172] In this step, the B side can obtain the calculated value P B of the second input variable according to the second input variable X B and the second model parameter W B = W BX B , and send P B to the server at end A.
[0173] Step S303: The third data terminal obtains the predicted value of the target variable based on the calculated value of the second input variable, the first input variable, and the first model parameter.
[0174] In this step, end A can calculate P based on the calculated value of the second input variable B , the first input variable X A , and the first model parameter W A to obtain the predicted value of the target variable Y′ = W A X A + P B .
[0175] Step S304: The third data terminal obtains the sum of squared residuals based on the dependent variable and the predicted value of the target variable, and obtains the sum of squared deviations based on the dependent variable and the mean value of the dependent variable.
[0176] In this step, end A can obtain the sum of squared residuals based on the dependent variable Y i , the predicted value of the target variable Y′, and obtain the sum of squared deviations based on Y i and the mean value of the dependent variable .
[0177] Step S305: The third data terminal determines the second correlation coefficient corresponding to the target variable based on the sum of squared residuals and the sum of squared deviations, and outputs the second correlation coefficient.
[0178] In this step, the second correlation coefficient is also called the variance inflation factor
[0179] As an optional embodiment, iteratively execute the step of selecting any feature variable in the first set of feature variables as the target variable until the second correlation coefficient corresponding to each feature variable when it is used as the target variable is output.
[0180] Specifically, repeat steps S301 - S305 a total of f_dim A times until the second correlation coefficients corresponding to all feature variables in the first set of feature variables in end A are output; similarly, any feature variable in the second set of feature variables can be iteratively selected as the target variable, that is, end A and end B exchange roles and execute the steps in the above embodiment, a total of f_dim B times, and the second correlation coefficients corresponding to all feature variables in the second set of feature variables in end B are output. In summary, a total of f_dim A + f_dim BAfter the secondary analysis, the second correlation coefficients corresponding to the feature variables of both parties will be output.
[0181] As an optional embodiment, after outputting the second correlation coefficient corresponding to each feature variable as the target variable, it further includes: selecting the feature variables whose second correlation coefficients meet the second preset condition to form a candidate data set.
[0182] Specifically, the second correlation coefficient, also known as the variance inflation factor VIF, can describe the degree of multicollinearity among various feature variables. In order to obtain a more reliable model, the feature variables with a relatively large variance inflation factor VIF can be removed. Figure 10 The following is a schematic diagram of the variance inflation factor VIF corresponding to each feature variable provided by the embodiment of the present invention. Figure 10 As shown in the figure, by observing the VIF values of all feature variables, if it is found that the VIF value is relatively large (significantly outlier), the feature variable can be removed, so as to obtain a feature combination with lower correlation to enhance the interpretability of the model.
[0183] The multi-variable processing method provided by the embodiment of the present invention is applied to the third data terminal in a multi-variable processing system. The multi-variable processing system includes the third data terminal and the fourth data terminal. Among them, the third data terminal stores a first set of feature variables, and the second data terminal stores a second set of feature variables; by selecting any feature variable in the first set of feature variables as the target variable, the other feature variables in the first set of feature variables as the first input variables, and the second set of feature variables as the second input variables; receiving the calculated values of the second input variables sent by the fourth data terminal, where the calculated values of the second input variables are obtained according to the second input variables and the second model parameters; obtaining the predicted value of the target variable according to the calculated values of the second input variables, the first input variables, and the first model parameters; obtaining the sum of squared residuals according to the target variable and the predicted value of the target variable, and obtaining the sum of squared deviations according to the target variable and the mean value of the target variable; determining the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and outputting the second correlation coefficient; that is, in the embodiment of the present invention, by transmitting intermediate parameters between different data owners, it is possible to analyze the linear correlation degree among various feature variables on the basis of ensuring the security of their respective data, providing an effective basis for subsequent feature screening.
[0184] On the basis of the above embodiments, Figure 11 The following is a schematic flowchart of a multi-variable processing method provided by Embodiment 5 of the present invention. Figure 11 As shown in the figure, the multi-variable processing method includes:
[0185] Step S401: Select any feature variable from the first set of feature variables as the target variable, the other feature variables in the first set of feature variables as the first input variables, and the second set of feature variables as the second input variables.
[0186] Step S402: The fourth data terminal constructs a second model based on the initial values of the second model parameters and the second input variables, encrypts the second model, and sends the encrypted second model to the third data terminal.
[0187] Correspondingly, on the side of the third data terminal, receive the encrypted second model sent by the fourth data terminal, where the second model is constructed based on the initial values of the second model parameters and the second input variables.
[0188] Step S403: The third data terminal constructs a first model based on the initial values of the first model parameters, the first input variables, and the target variable.
[0189] Optionally, it further includes steps S404, S406, S408, S410, and S412. That is, on the side of the third data terminal, encrypt the first model and send the encrypted first model to the fourth data terminal. The encrypted first model is used by the fourth data terminal to calculate the encrypted second gradient value of the global loss function with respect to the second model parameters, and add the encrypted second gradient value to the second random number to obtain the encrypted second gradient-related value; receive the encrypted second gradient-related value sent by the fourth data terminal, decrypt the encrypted second gradient-related value, and send the decrypted second gradient-related value to the fourth data terminal. The decrypted second gradient-related value is used by the fourth data terminal to obtain the second gradient value and update the second model parameters according to the second gradient value. Correspondingly, on the side of the fourth data terminal, receive the encrypted first model sent by the third data terminal, where the first model is constructed based on the initial values of the first model parameters, the first input variables, and the target variable; calculate the encrypted second gradient value of the global loss function with respect to the second model parameters according to the encrypted first model, the second model, and the second input variables; add the encrypted second gradient value and the second random number to obtain the encrypted second gradient-related value, and send the encrypted second gradient-related value to the third data terminal for decryption; receive the decrypted second gradient-related value sent by the third data terminal, and obtain the second gradient value according to the second gradient-related value; update the second model parameters according to the second gradient value.
[0190] Step S404: The third data terminal encrypts the first model and sends the encrypted first model to the fourth data terminal.
[0191] Step S405: The third data terminal calculates the encrypted first gradient value of the global loss function with respect to the first model parameters based on the encrypted second model, the first model, and the first input variable.
[0192] Step S406: The fourth data terminal calculates the encrypted second gradient value of the global loss function with respect to the second model parameters based on the encrypted first model, the second model, and the second input variable.
[0193] Step S407: The third data terminal adds the encrypted first gradient value and the first random number to obtain an encrypted first gradient-related value, and sends the encrypted first gradient-related value to the fourth data terminal.
[0194] Step S408: The fourth data terminal adds the encrypted second gradient value and the second random number to obtain an encrypted second gradient-related value, and sends the encrypted second gradient-related value to the third data terminal.
[0195] Step S409: The fourth data terminal receives the encrypted first gradient-related value, decrypts the encrypted first gradient-related value, and sends the decrypted first gradient-related value to the third data terminal.
[0196] Step S410: The third data terminal receives the encrypted second gradient-related value sent by the fourth data terminal, decrypts the encrypted second gradient-related value, and sends the decrypted second gradient-related value to the fourth data terminal.
[0197] Step S411: The third data terminal receives the decrypted first gradient-related value sent by the fourth data terminal, obtains the first gradient value based on the first gradient-related value, and updates the first model parameters according to the first gradient value.
[0198] Step S412: The fourth data terminal receives the decrypted second gradient-related value sent by the third data terminal, obtains the second gradient value based on the second gradient-related value, and updates the second model parameters according to the second gradient value.
[0199] Iteratively execute steps S402 - S413 until the regression model constructed by the first input variable, the second input variable, and the target variable converges and the global loss function reaches the minimum value.
[0200] Step S413: The fourth data terminal obtains the calculated value of the second input variable based on the second input variable and the second model parameters, and sends the calculated value of the second input variable to the third data terminal.
[0201] Step S414: The third data terminal obtains the predicted value of the target variable based on the calculated value of the second input variable, the first input variable, and the first model parameters.
[0202] Step S415: The third data terminal obtains the sum of squared residuals based on the dependent variable and the predicted value of the target variable, and obtains the sum of squared deviations based on the dependent variable and the mean value of the dependent variable.
[0203] Step S416: The third data terminal determines the second correlation coefficient between the dependent variable and the first input variable and the second input variable according to the sum of squared residuals and the sum of squared deviations, and outputs the second correlation coefficient.
[0204] The implementation manners of step S401, step S413-step S416 in this embodiment are respectively similar to those of step S301-step S305 in the above embodiment, and will not be elaborated here.
[0205] The difference from the above embodiment is that this embodiment further defines how to determine the model parameters constructed by the first input variable, the second input variable and the target variable. In this embodiment, the following steps are iteratively executed until the regression model constructed by the first input variable, the second input variable and the target variable converges and the global loss function obtains the minimum value: receiving the encrypted second model sent by the fourth data terminal, where the second model is constructed according to the initial value of the second model parameters and the second input variable; constructing a first model according to the initial value of the first model parameters, the first input variable and the target variable; calculating the encrypted first gradient value of the global loss function with respect to the first model parameters according to the encrypted second model, the first model and the first input variable; adding the encrypted first gradient value and the first random number to obtain an encrypted first gradient related value, and sending the encrypted first gradient related value to the fourth data terminal for decryption; receiving the decrypted first gradient related value sent by the fourth data terminal, and obtaining the first gradient value according to the first gradient related value; updating the first model parameters according to the first gradient value; further including: the third data terminal encrypts the first model and sends the encrypted first model to the fourth data terminal, and the encrypted first model is used for the fourth data terminal to calculate the encrypted second gradient value of the global loss function with respect to the second model parameters, and adding the encrypted second gradient value and the second random number to obtain an encrypted second gradient related value; receiving the encrypted second gradient related value sent by the fourth data terminal, decrypting the encrypted second gradient related value, and sending the decrypted second gradient related value to the fourth data terminal, where the decrypted second gradient related value is used for the fourth data terminal to obtain the second gradient value and update the second model parameters according to the second gradient value.
[0206] Specifically, the embodiment of the present invention performs regression on the linear model by the gradient descent method until the regression model constructed by the first input variable, the second input variable and the target variable converges and the global loss function obtains the minimum value.
[0207] For the first model at end A, first initialize the first model parameters at end A to obtain the initial value w of the first model parameters A , initialize the second model parameters at end B to obtain the initial value w of the second model parameters B ; then end B constructs the second model F B based on w B and the second input variable X B = w B X B , encrypts F B , and sends the encrypted F B to end A.
[0208] Optionally, the fourth data end includes a fourth key pair, and the fourth key pair includes a fourth public key and a fourth private key; among them, the encrypted second model is obtained by encrypting with the fourth public key. Specifically, the fourth key pair generated by end B includes the fourth public key PK B and the fourth private key SK B , end B uses PK B to encrypt F B to obtain the encrypted second model Enc B (F B ), and sends Enc B (F B ) to end A, and end A cannot decrypt it.
[0209] Then, end A constructs the first model F A based on the initial value w of the first model parameters A , the first input variable X A , and the target variable Y, where F A = w A X A - Y.
[0210] Then, end A calculates the encrypted first gradient value of the global loss function with respect to the first model parameters Enc B (G B ) based on the encrypted second model Enc A (F A ), the first model F B (G A ) = (F A + Enc B (F B ))X A .
[0211] Then, end A sends the encrypted first gradient value Enc B (G A ) and the first random number RA Addition processing, where R A is a vector composed of random values and has the same dimension as G A to obtain the encrypted first gradient-related value Enc B (G A +R A ); Send Enc B (G A +R A ) to Party B, and Party B uses the fourth private key SK B to decrypt Enc B (G A +R A ) and send G A +R A to Party A. Party A subtracts R A to obtain G A , and then uses G A to update the first model parameter.
[0212] Similarly, for the second model of Party B, after Party A obtains the first model F A =w A X A -Y, it will encrypt F A and send the encrypted F A to Party B.
[0213] Optionally, the method further includes: a third data terminal generates a third key pair, and the third key pair includes a third public key and a third private key; the encrypting process of the first model includes: encrypting the first model through the third public key. Specifically, the third key pair generated by Party A includes a third public key PK A and a fourth private key SK A . Party A uses PK A to encrypt F A to obtain Enc A (F A ).
[0214] Then, Party B calculates the encrypted second gradient value Enc A (G A ) of the global loss function with respect to the second model parameter according to Enc B (F B ), F A (G B )=(Enc A (F A )+F B )X B ; Then Party B sends Enc A (GB ) and the second random number R B for addition processing, where R B is a vector composed of random values and has the same dimension as G B , to obtain the encrypted second gradient-related value Enc A (G B + R B ); send Enc A (G B + R B ) to end A, and end A uses the third private key SK A to decrypt Enc A (G B + R B ), send G B + R B to end B, and end B subtracts R B to obtain G B , and then uses G B to update the second model parameter.
[0215] As an optional embodiment, the method further includes: the third data end calculates the first local loss function according to the first model, calculates the encrypted third local loss function according to the first model and the encrypted second model, and receives the encrypted second local loss function calculated by the fourth data end according to the second model; obtains the encrypted global loss function value according to the first local loss function, the encrypted second local loss function and the third local loss function, and sends the encrypted global loss function value to the fourth data end for decryption; receives the decrypted global loss function value and the regression model convergence flag sent by the fourth data end.
[0216] Specifically, the global loss function is shown in formula (12):
[0217]
[0218] Decompose the global loss function Loss, and the global loss function Loss can be divided into the local loss function F sqrA related to the first model, the local loss function F sqrB related to the second model, and the local loss function including the first model and the second model, as shown in formula (13):
[0219]
[0220] Therefore, in order to obtain the global loss function, end A calculates the first local loss function F A according to the first model F sqrA , according to the F A , the encrypted second model EncB (F B ) Calculate to obtain the encrypted third partial loss function, and receive the encrypted second partial loss function F calculated by Party B according to the second model F B Calculate to obtain the encrypted second partial loss function F sqrB ; Obtain the value of the encrypted global loss function Loss according to the first partial loss function, the encrypted second partial loss function, and the third partial loss function, and send the value of the encrypted global loss function Loss to Party B for decryption; Receive the decrypted global loss function value and the regression model convergence flag sent by Party B. When the regression model converges and the global loss function obtains the minimum value, determine the updated first model parameter and second model parameter as the final model parameter for calculating the second correlation coefficient of each independent variable.
[0221] The multi-variable processing method provided by the embodiment of the present invention iteratively executes the following steps until the regression model constructed by the first input variable, the second input variable, and the target variable converges and the global loss function obtains the minimum value: Receive the encrypted second model sent by the fourth data terminal, where the second model is constructed according to the initial value of the second model parameter and the second input variable; Construct the first model according to the initial value of the first model parameter, the first input variable, and the target variable; Calculate to obtain the encrypted first gradient value of the global loss function with respect to the first model parameter according to the encrypted second model, the first model, and the first input variable; Add the encrypted first gradient value and the first random number to obtain the encrypted first gradient correlation value, and send the encrypted first gradient correlation value to the fourth data terminal for decryption; Receive the decrypted first gradient correlation value sent by the fourth data terminal, and obtain the first gradient value according to the first gradient correlation value; Update the first model parameter according to the first gradient value; It further includes: The third data terminal encrypts the first model and sends the encrypted first model to the fourth data terminal, and the encrypted first model is used for the fourth data terminal to calculate the encrypted second gradient value of the global loss function with respect to the second model parameter, and add the encrypted second gradient value and the second random number to obtain the encrypted second gradient correlation value; Receive the encrypted second gradient correlation value sent by the fourth data terminal, decrypt the encrypted second gradient correlation value, and send the decrypted second gradient correlation value to the fourth data terminal. The decrypted second gradient correlation value is used for the fourth data terminal to obtain the second gradient value and update the second model parameter according to the second gradient value; That is, the embodiment of the present invention uses the gradient descent method to perform regression on each model to obtain more stable model parameters.
[0222] To further understand the embodiment of the present invention, Figure 12 is a schematic flowchart of a multi-variable processing method provided by Embodiment VI of the present invention, combined with Figure 1 and Figure 12, the multi-variable processing method includes:
[0223] Step 1: B generates a pair of public and private keys PK B , SK B , and sends the public key PK B to A;
[0224] Step 2: A generates a pair of public and private keys PK A , SK A (PK A ≠ PK B , SK A ≠ SK B ), and sends the public key PK A to B;
[0225] Step 3: B sends its own feature dimension f_dim B to A;
[0226] Step 4: A sends its own feature dimension f_dim A to B;
[0227] Step 5: A extracts from the full data
[0228] Step 6: B calculates F B = w B X B , encrypts F B with its own public key PK B , and sends Enc B (F B ) to A;
[0229] Step 7: A calculates F A = w A X A - Y, encrypts F A with its own public key PK A , and sends Enc A (F A ) to B;
[0230] Step 8: B calculates Enc A (G B ) = (Enc A (F A ) + F B )X B , and sends Enc A (G B + R B ) to A, where R B is a vector composed of random values (with the same dimension as G B );
[0231] Step 9, A calculates Enc B (G A ) = (F A + Enc B (F B ))X A , and sends Enc B (G A + R A ) to B, where R A is a vector composed of random values (with the same dimension as G A );
[0232] Step 10, B decrypts Enc B (G B + R A + R A ) using the private key SK A + R A to A;
[0233] Step 11, A decrypts Enc A (G A + R B + R B ) using the private key SK B + R B to B;
[0234] Step 12, B calculates and sends Enc B (F sqrB ) to A;
[0235] Step 13, A calculates and sends Enc B (L) + Enc B (L normA ) to B;
[0236] Step 14. B decrypts Enc B (L) + Enc B (L normA ), calculates L total = L + L normA + L normB , and determines whether the model is currently fitting, and sends L total and the fitting flag (true / false) to A;
[0237] Among them, L normA is the regularization term at the A end, and L normB is the regularization term at the B end, which is used to limit the overfitting of the model.
[0238] Step 15. Each of the two parties performs gradient optimization locally and updates the model weights;
[0239] Step 16. Iteratively perform Step 5 to Step 15 until the model effect meets the requirements;
[0240] Step 17. Party B calculates P B = W B X B , and sends P B to Party A;
[0241] Step 18. Party A calculates Y' = W A X A + P B , (If the current Y is the feature of Party B, Party B needs to send VIF A to Party A after calculation);
[0242] Step 19. Repeat Step 5 to Step 8 until all feature variables have been analyzed, that is, first analyze all the features of Party A in sequence, and then analyze all the features of Party B in sequence. A total of (f_dim A + f_dim B ) times of analysis are required.
[0243] Among them, Step 1 and Step 2 are to ensure that the information transmitted between the two parties cannot be cracked by a third party. Then the server at Party A's end can encrypt the message sent to the server at Party B's end using the private key SK A , and the server at Party B's end can decrypt the message using the public key PK A received from the server at Party A's end; similarly, the server at Party B's end can encrypt the message sent to the server at Party A's end using SK B , and the server at Party A's end can decrypt the message using the public key PK B received from the server at Party B's end.
[0244] In summary, the server at Party A's end and the server at Party B's end respectively have different data. The two parties can perform data-related calculations locally. Only the encrypted intermediate parameters are transmitted between the two parties, and the receiving party cannot reverse-engineer the original data. On the basis of ensuring the security of their respective data, the degree of multicollinearity corresponding to each independent variable is analyzed, providing an effective basis for subsequent variable screening, etc.
[0245] The embodiment of the present invention also provides a third data terminal. Figure 13 It is a schematic structural diagram of a third data terminal provided by the embodiment of the present invention. As Figure 13 shown, the third data terminal includes a third processing module 30 and a third receiving module 31;
[0246] Among them, the third processing module 30 is configured to select any feature variable in the first feature variable set as the target variable, other feature variables in the first feature variable set as the first input variables, and the second feature variable set as the second input variables; a third receiving module 31 is configured to receive the calculated value of the second input variables sent by the fourth data terminal, where the calculated value of the second input variables is obtained according to the second input variables and the second model parameters; the third processing module 30 is further configured to obtain a predicted value of the target variable according to the calculated value of the second input variables, the first input variables, and the first model parameters; obtain the sum of squared residuals according to the target variable and the predicted value of the target variable, and obtain the sum of squared deviations according to the target variable and the mean value of the target variable; determine the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and output the second correlation coefficient.
[0247] As an optional implementation manner, the multi-variable processing system further includes a third sending module 32; the third processing module 30, the third receiving module 31, and the third sending module 32 are further configured to: iteratively execute the following steps until the regression model constructed by the first input variables, the second input variables, and the target variable converges and the global loss function obtains the minimum value: the third receiving module 31 receives the encrypted second model sent by the fourth data terminal, where the second model is obtained according to the initial values of the second model parameters and the second input variables; the third processing module 30 constructs a first model according to the initial values of the first model parameters, the first input variables, and the target variable; calculate the encrypted first gradient value of the global loss function with respect to the first model parameters according to the encrypted second model, the first model, and the first input variables; perform an addition process on the encrypted first gradient value and a first random number to obtain an encrypted first gradient-related value, and send the encrypted first gradient-related value to the fourth data terminal through the third sending module 32 for decryption; the third receiving module 31 receives the decrypted first gradient-related value sent by the fourth data terminal, and obtains the first gradient value according to the first gradient-related value through the third processing module 30; update the first model parameters according to the first gradient value.
[0248] As an optional implementation manner, the fourth data terminal includes a fourth key pair, where the fourth key pair includes a fourth public key and a fourth private key; among them, the encrypted second model is obtained by encrypting with the fourth public key; the decrypted first gradient-related value is obtained by decrypting with the fourth private key.
[0249] As an alternative implementation, the third processing module 30 is further configured to encrypt the first model and send the encrypted first model to the fourth data terminal. The encrypted first model is used for the fourth data terminal to calculate the encrypted second gradient value of the global loss function with respect to the second model parameters, and add the encrypted second gradient value to a second random number to obtain an encrypted second gradient-related value. The third receiving module 31 is further configured to receive the encrypted second gradient-related value sent by the fourth data terminal, decrypt the encrypted second gradient-related value, and send the decrypted second gradient-related value to the fourth data terminal. The decrypted second gradient-related value is used for the fourth data terminal to obtain the second gradient value and update the second model parameters according to the second gradient value.
[0250] As an alternative implementation, the third processing module 30 is further configured to: generate a third key pair, where the third key pair includes a third public key and a third private key; encrypt the first model with the third public key; and decrypt the encrypted second gradient-related value with the third private key.
[0251] As an alternative implementation, the third processing module 30 is further configured to calculate a first local loss function according to the first model, calculate an encrypted third local loss function according to the first model and the encrypted second model, and receive, through the third receiving module 31, the encrypted second local loss function calculated by the fourth data terminal according to the second model. The third processing module 30 obtains an encrypted global loss function value according to the first local loss function, the encrypted second local loss function, and the third local loss function, and sends the encrypted global loss function value to the fourth data terminal through the third sending module 32 for decryption. The third receiving module 31 receives the decrypted global loss function value and the regression model convergence flag sent by the fourth data terminal.
[0252] As an alternative implementation, the third processing module 30 is configured to iteratively execute the step of selecting any feature variable in the first feature variable set as the target variable until the second correlation coefficient corresponding to each feature variable as the target variable is output.
[0253] As an alternative implementation, the third processing module 30 is configured to select the feature variables whose second correlation coefficients meet the second preset condition to form a candidate data set.
[0254] The implementation principle and technical effects of the third data terminal provided in this embodiment are similar to those of the above embodiments, and will not be elaborated here.
[0255] An embodiment of the present invention further provides a fourth data terminal. Figure 14 The structural schematic diagram of a fourth data terminal provided by an embodiment of the present invention is asFigure 14 As shown, the fourth data terminal includes a fourth processing module 40 and a fourth sending module 41; wherein, the fourth processing module 40 is configured to obtain a calculated value of the second input variable according to the second input variable and the second model parameter; the fourth sending module 41 is configured to send the calculated value of the second input variable to the third data terminal, wherein the calculated value of the second input variable is used to determine a second correlation coefficient corresponding to the target variable.
[0256] As an optional embodiment, the fourth data terminal further includes a fourth receiving module 42; the fourth processing module 40 is further configured to iteratively execute the following steps until the regression model constructed by the first input variable, the second input variable, and the target variable converges and the global loss function obtains a minimum value: construct a second model according to the initial value of the second model parameter and the second input variable, encrypt the second model, and send the encrypted second model to the third data terminal through the fourth sending module 41, where the encrypted second model is used for the third data terminal to obtain an encrypted first gradient value of the global loss function with respect to the first model parameter, and add the encrypted first gradient value and a first random number to obtain an encrypted first gradient-related value; the fourth receiving module 42 is configured to receive the encrypted first gradient-related value sent by the third data terminal, decrypt the encrypted first gradient-related value through the fourth processing module 40, and send the decrypted first gradient-related value to the third data terminal through the fourth invention module 41, where the decrypted first gradient-related value is used for the third data terminal to obtain a first gradient value and update the first model parameter according to the first gradient value.
[0257] As an optional embodiment, the fourth receiving module 41 is further configured to receive an encrypted first model sent by the third data terminal, where the first model is constructed according to the initial value of the first model parameter, the first input variable, and the target variable; the fourth processing module 40 is further configured to calculate an encrypted second gradient value of the global loss function with respect to the second model parameter according to the encrypted first model, the second model, and the second input variable; add the encrypted second gradient value and a second random number to obtain an encrypted second gradient-related value, and send the encrypted second gradient-related value to the third data terminal for decryption; the fourth receiving module 42 is configured to receive the decrypted second gradient-related value sent by the third data terminal, and obtain a second gradient value according to the second gradient-related value through the fourth processing module 40; update the second model parameter according to the second gradient value.
[0258] The implementation principle and technical effect of the fourth data terminal provided in this embodiment are similar to those of the above embodiment, and will not be elaborated here.
[0259] The embodiment of the present invention further provides a multi-variable processing system. Reference may be made to Figure 1As shown, the multivariable system includes a third data terminal and a fourth data terminal; wherein, the third data terminal and the fourth data terminal cooperate to execute the method according to any one of the embodiments.
[0260] The multivariable processing system provided in this embodiment has the same implementation principle and technical effects as the above embodiments, and will not be elaborated here.
[0261] The present invention also provides a variable screening method, which is applied to a variable screening system. The variable screening system includes a fifth data terminal and a sixth data terminal. Among them, the fifth data terminal stores a dependent variable and a first independent variable set, and the sixth data terminal stores a second independent variable set. For reference, Figure 1 as shown ( Figure 1 the A-side server in it is equivalent to the fifth data terminal, and the B-side server is equivalent to the sixth data terminal).
[0262] The variable screening method includes steps S501 - S517:
[0263] Step S501, the fifth data terminal obtains the first correlation coefficient between each independent variable in the first independent variable set and the dependent variable.
[0264] Specifically, the A-side stores the dependent variable and some independent variables (i.e., the first independent variable set) corresponding to the sample data, and the B-side stores the other part of the independent variables (i.e., the second independent variable set) corresponding to the sample data. First, the A-side obtains the linear correlation degree between the independent variables on the A-side and the dependent variable (i.e., the first correlation coefficient, the variance inflation coefficient R in the above embodiments 2 ).
[0265] Step S502, the fifth data terminal obtains the difference between the dependent variable and the mean value of the dependent variable, and sends the difference to the sixth data terminal.
[0266] Step S503, the sixth data terminal receives the difference, and calculates the regression coefficient of the unary linear regression model constructed by any independent variable in the first independent variable set and the dependent variable according to the difference.
[0267] Step S504, the sixth data terminal obtains a third parameter according to the regression coefficient and the mean value of the independent variable, obtains a fourth parameter according to the regression coefficient and the independent variable, encrypts the fourth parameter, and sends the third parameter and the encrypted fourth parameter to the first data terminal.
[0268] Step S505: The fifth data terminal receives the third parameter and the encrypted fourth parameter, obtains the constant term of the unary linear regression model according to the third parameter and the mean value of the dependent variable, obtains the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term, and the dependent variable, and obtains the sum of squared deviations according to the difference between the dependent variable and the mean value of the dependent variable; obtains the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and sends the encrypted first correlation coefficient to the sixth data terminal.
[0269] Step S506: The sixth data terminal receives the encrypted first correlation coefficient, decrypts the encrypted first correlation coefficient, and sends the decrypted first correlation coefficient to the fifth data terminal.
[0270] Step S507: The fifth data terminal receives the decrypted first correlation coefficient and outputs the first correlation coefficient; iteratively execute the step of obtaining the difference between the dependent variable and the mean value of the dependent variable until the first correlation coefficient corresponding to each independent variable in the second independent variable is output.
[0271] Step S508: Select the independent variables in the first independent variable set whose first correlation coefficients meet the first preset condition to form a first candidate data set, and select the independent variables in the second independent variable set whose first correlation coefficients meet the first preset condition to form a second candidate data set.
[0272] Specifically, the implementation manners of steps S502 - S507 are similar to the implementation manner of the single-variable processing method described in the first aspect, and are used to obtain the variance inflation factor R of each independent variable at the B end and the dependent variable at the A end 2 , and then select the independent variables in the first independent variable set whose variance inflation factor R 2 is greater than a certain threshold (for example, 0.3) to form a first candidate data set, and select the independent variables in the second independent variable set whose variance inflation factor R 2 is greater than a certain threshold (for example, 0.3) to form a second candidate data set.
[0273] Step S509: Select any independent variable in the first candidate data set as the target variable, the other independent variables in the first candidate data set as the first input variables, and the second candidate data set as the second input variables.
[0274] Step S510: The sixth data terminal obtains the calculated values of the second input variables according to the second input variables and the second model parameters, and sends the calculated values of the second input variables to the fifth data terminal.
[0275] Step S511: The fifth data terminal obtains the predicted value of the target variable based on the calculated value of the second input variable, the first input variable, and the first model parameter; obtains the sum of squared residuals based on the target variable and the predicted value of the target variable, and obtains the sum of squared deviations based on the target variable and the mean value of the target variable; determines the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and outputs the second correlation coefficient.
[0276] Step S512: Iteratively execute the step of selecting any independent variable in the first candidate data set as the target variable until the second correlation coefficients corresponding to each independent variable in the first candidate data set when it is used as the target variable are output.
[0277] Step S513: Select any independent variable in the second candidate data set as the target variable, other independent variables in the second candidate data set as the third input variable, and the first candidate data set as the fourth input variable.
[0278] Step S514: The fifth data terminal obtains the calculated value of the fourth input variable based on the fourth input variable and the fourth model parameter, and sends the calculated value of the fourth input variable to the sixth data terminal.
[0279] Step S515: The sixth data terminal obtains the predicted value of the target variable based on the calculated value of the fourth input variable, the third input variable, and the third model parameter; obtains the sum of squared residuals based on the target variable and the predicted value of the target variable, and obtains the sum of squared deviations based on the target variable and the mean value of the target variable; determines the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and outputs the second correlation coefficient.
[0280] Step S516: Iteratively execute the step of selecting any independent variable in the second candidate data set as the target variable until the second correlation coefficients corresponding to each independent variable in the second candidate data set when it is used as the target variable are output.
[0281] Step S517: Select the independent variables in the first candidate data set whose second correlation coefficients meet the second preset condition to form the third candidate data set, and select the independent variables in the second candidate data set whose second correlation coefficients meet the second preset condition to form the fourth candidate data set. The third candidate data set and the fourth candidate data set form the final candidate data set.
[0282] Specifically, after obtaining the first candidate data set and the second candidate data set, then execute the multivariate processing method as described in the above embodiment to obtain the degree of multicollinearity, i.e., the variance inflation factor VIF, between each independent variable in the two candidate data sets and other independent variables, and eliminate the independent variables with larger VIF values to form the final candidate data set.
[0283] In summary, for the variable screening method provided in the embodiments of the present invention, first, by analyzing the variance inflation factor R of each independent variable and the dependent variable 2 , the independent variables with larger R 2 are obtained. Then, by analyzing the degree of multicollinearity VIF of each independent variable, the feature combinations with larger VIF values are selected, providing an effective basis for subsequent feature screening.
[0284] Another variable screening provided in the embodiments of the present invention includes the following steps:
[0285] Step S601: Select any independent variable in the first set of independent variables as the target variable, the other independent variables in the first set of independent variables as the first input variables, and the second set of independent variables as the second input variables.
[0286] Step S602: The sixth data terminal obtains the calculated value of the second input variable according to the second input variable and the second model parameter, and sends the calculated value of the second input variable to the fifth data terminal.
[0287] Step S603: The fifth data terminal obtains the predicted value of the target variable according to the calculated value of the second input variable, the first input variables, and the first model parameter; obtains the sum of squared residuals according to the target variable and the predicted value of the target variable, and obtains the sum of squared deviations according to the target variable and the mean value of the target variable; determines the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and outputs the second correlation coefficient.
[0288] Step S604: Iteratively execute the step of selecting any independent variable in the first set of independent variables as the target variable until the second correlation coefficients corresponding to each independent variable in the first set of independent variables are output when used as the target variable.
[0289] Step S605: Select any independent variable in the second set of independent variables as the target variable, the other independent variables in the second set of independent variables as the third input variables, and the second set of independent variables as the fourth input variables.
[0290] Step S606: The fifth data terminal obtains the calculated value of the fourth input variable according to the fourth input variable and the fourth model parameter, and sends the calculated value of the fourth input variable to the sixth data terminal.
[0291] Step S607: The sixth data terminal obtains the predicted value of the target variable according to the calculated value of the fourth input variable, the third input variables, and the third model parameter; obtains the sum of squared residuals according to the target variable and the predicted value of the target variable, and obtains the sum of squared deviations according to the target variable and the mean value of the target variable; determines the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and outputs the second correlation coefficient.
[0292] Step S608: Iteratively execute the step of selecting any independent variable in the second set of independent variables as the target variable until the second correlation coefficients corresponding to each independent variable in the second set of independent variables when used as the target variable are output.
[0293] Step S609: Select the independent variables in the first set of independent variables whose second correlation coefficients meet the second preset condition to form a fifth candidate data set, and select the independent variables in the second set of independent variables whose second correlation coefficients meet the second preset condition to form a sixth candidate data set.
[0294] Step S610: The fifth data terminal obtains the first correlation coefficients between each independent variable and the dependent variable in the fifth candidate data set.
[0295] Step S611: The fifth data terminal obtains the difference between the dependent variable and the mean value of the dependent variable, and sends the difference to the sixth data terminal.
[0296] Step S612: The sixth data terminal receives the difference, and calculates the regression coefficient of the simple linear regression model constructed by any independent variable and the dependent variable in the sixth candidate data set based on the difference.
[0297] Step S613: The sixth data terminal obtains a third parameter based on the regression coefficient and the mean value of the independent variable, obtains a fourth parameter based on the regression coefficient and the independent variable, performs encryption processing on the fourth parameter, and sends the third parameter and the encrypted fourth parameter to the fifth data terminal.
[0298] Step S614: The fifth data terminal receives the third parameter and the encrypted fourth parameter, obtains the constant term of the simple linear regression model based on the third parameter and the mean value of the dependent variable, obtains the encrypted residual sum of squares based on the encrypted fourth parameter, the constant term, and the dependent variable, and obtains the sum of squares of deviations based on the difference between the dependent variable and the mean value of the dependent variable; obtains the encrypted first correlation coefficient corresponding to the independent variable based on the encrypted residual sum of squares and the sum of squares of deviations, and sends the encrypted first correlation coefficient to the sixth data terminal.
[0299] Step S615: The sixth data terminal receives the encrypted first correlation coefficient, decrypts the encrypted first correlation coefficient, and sends the decrypted first correlation coefficient to the fifth data terminal.
[0300] Step S616: The fifth data terminal receives the decrypted first correlation coefficient and outputs the first correlation coefficient;
[0301] Step S617: Iteratively execute the step of obtaining the difference between the dependent variable and the mean value of the dependent variable until the first correlation coefficients corresponding to each independent variable in the sixth candidate data set are output.
[0302] Step S618: Select the independent variables in the fifth candidate data set whose first correlation coefficient meets the first preset condition to form a seventh candidate data set, and select the independent variables in the sixth candidate data set whose first correlation coefficient meets the first preset condition to form an eighth candidate data set. The seventh candidate data set and the eighth candidate data set form the final candidate data set.
[0303] In summary, for the variable screening method provided by the embodiments of the present invention, first, by analyzing the degree of multicollinearity VIF of each independent variable, a feature combination with a larger VIF value is selected; then, the variance inflation factor R of each independent variable in the feature combination with a larger VIF value and the dependent variable is analyzed 2 , and R 2 with a larger value is obtained, providing an effective basis for subsequent feature screening.
[0304] The present invention also provides a variable screening system, as Figure 1 shown. The variable screening system includes a fifth data terminal and a sixth data terminal; wherein, the fifth data terminal and the sixth data terminal are used to implement the above-mentioned variable screening method.
[0305] As Figure 15 shown, the embodiments of the present invention provide an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114,
[0306] The memory 113 is used to store a computer program;
[0307] In an embodiment of the present invention, when the processor 111 is used to execute the program stored on the memory 113, it implements the steps of the single-variable processing method or the multi-variable processing method provided by any one of the foregoing method embodiments.
[0308] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the single-variable processing method or the multi-variable processing method provided by any one of the foregoing method embodiments.
[0309] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0310] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather should be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A single variable processing method, characterized in that, Applied to the first data end in a univariate processing system, the univariate processing system includes the first data end and the second data end. Among them, the first data end stores a dependent variable and a first set of characteristic variables, and the second data end stores a second set of characteristic variables, and the second set of characteristic variables includes independent variables; the method includes: Obtain the difference between the dependent variable and the mean value of the dependent variable, and send the difference to the second data end. The difference is used to calculate the regression coefficient of the univariate linear regression model constructed by the independent variable and the dependent variable; Receive the third parameter and the encrypted fourth parameter sent by the second data end. Among them, the third parameter is calculated based on the regression coefficient and the mean value of the independent variable, and the fourth parameter is calculated based on the regression coefficient and the independent variable; Obtain the constant term of the univariate linear regression model according to the third parameter and the mean value of the dependent variable; Obtain the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term and the dependent variable, and obtain the sum of squared deviations according to the difference between the dependent variable and the mean value of the dependent variable; Obtain the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and send the encrypted first correlation coefficient to the second data end for decryption; Receive the decrypted first correlation coefficient sent by the second data end and output the first correlation coefficient; The method further includes: Obtain the calculated value of the second input variable according to the second input variable and the second model parameter, and send the calculated value of the second input variable to the second data end; Among them, the second data end selects the independent variable in the second set of characteristic variables as the target variable, other characteristic variables in the second set of characteristic variables as the first input variable, the first set of characteristic variables as the second input variable. The second data end obtains the predicted value of the target variable according to the calculated value of the second input variable, the first input variable and the first model parameter, obtains the sum of squared residuals according to the dependent variable and the predicted value of the target variable, and obtains the sum of squared deviations according to the dependent variable and the mean value of the dependent variable. Determine the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and output the second correlation coefficient.
2. The method according to claim 1, wherein The obtaining the difference between the dependent variable and the mean value of the dependent variable, and sending the difference to the second data end, where the difference is used to calculate the regression coefficient of the univariate linear regression model constructed by the independent variable and the dependent variable, includes: Receive the encrypted first parameter sent by the second data end, and the first parameter is calculated based on the independent variable and the mean value of the independent variable; Obtain the encrypted second parameter according to the encrypted first parameter and the difference between the dependent variable and the mean value of the dependent variable, and send the encrypted second parameter to the second data end for decryption. The decrypted second parameter is used to calculate the regression coefficient of the univariate linear regression model constructed by the independent variable and the dependent variable.
3. The method according to claim 2, wherein The second data end includes a second key pair, and the second key pair includes a second public key and a second private key; Among them, the encrypted fourth parameter and the encrypted first parameter are both encrypted by the second public key; The decrypted first correlation coefficient and the decrypted second parameter are both obtained by decrypting with the second private key.
4. The method according to any one of claims 1 to 3, characterized in that, The independent variables stored at the second data end are processed by binning; before obtaining the difference between the dependent variable and the mean value of the dependent variable, it further includes: Sending the sample identifier and the encrypted sample tag value corresponding to the sample identifier to the second data end; Receiving the encrypted sample tag statistical value sent by the second data end, where the encrypted sample tag statistical value is obtained by the second data end statistically processing the encrypted sample tag values of each bin according to the sample identifier; Performing decryption processing on the encrypted sample tag statistical value to obtain the dependent variable.
5. The method according to claim 4, wherein Before sending the sample identifier and the encrypted sample tag value corresponding to the sample identifier to the second data end, it further includes: Generating a first key pair, where the first key pair includes a first public key and a first private key; Encrypting the sample tag value with the first public key to obtain the encrypted sample tag value; The performing decryption processing on the encrypted sample tag statistical value includes: Performing decryption processing on the encrypted sample tag statistical value with the first private key.
6. The method according to any one of claims 1-3, characterized in that, If the second data end includes multiple independent variables, iteratively execute the step of obtaining the difference between the dependent variable and the mean value of the dependent variable until the first correlation coefficient corresponding to each independent variable is output.
7. The method according to claim 6, characterized in that After outputting the first correlation coefficient corresponding to each independent variable, it further includes: Selecting the independent variables whose first correlation coefficients meet the first preset condition to form a candidate data set.
8. A single-variable processing method, characterized in that, Applied to the second data end in a univariate processing system, the univariate processing system includes a first data end and the second data end, where the first data end stores a dependent variable and a first feature variable set, the second data end stores a second feature variable set, and the second feature variable set includes independent variables; the method includes: Receiving the difference between the dependent variable and the mean value of the dependent variable sent by the first data end, and calculating the regression coefficient of the unary linear regression model constructed by the independent variable and the dependent variable according to the difference; Obtaining a third parameter according to the regression coefficient and the mean value of the independent variable, obtaining a fourth parameter according to the regression coefficient and the independent variable, and performing encryption processing on the fourth parameter; Sending the third parameter and the encrypted fourth parameter to the first data end, where the third parameter and the encrypted fourth parameter are used to calculate the encrypted first correlation coefficient corresponding to the independent variable; Receiving the encrypted first correlation coefficient sent by the first data end, decrypting the encrypted first correlation coefficient, sending the decrypted first correlation coefficient to the first data end, and outputting the first correlation coefficient; The method further includes: selecting the independent variables in the second feature variable set as the target variables, other feature variables in the second feature variable set as the first input variables, and the first feature variable set as the second input variables; Receiving the calculated value of the second input variable sent by the first data end, where the calculated value of the second input variable is obtained according to the second input variable and the second model parameter; Obtaining the predicted value of the target variable according to the calculated value of the second input variable, the first input variable, and the first model parameter; Obtain the sum of squared residuals based on the predicted values of the dependent variable and the target variable, and obtain the sum of squared deviations based on the dependent variable and the mean of the dependent variable; Determine the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and output the second correlation coefficient.
9. The method according to claim 8, wherein Receiving the difference between the dependent variable and the mean of the dependent variable sent by the first data terminal, and calculating the regression coefficient of the simple linear regression model constructed by the independent variable and the dependent variable, including: Calculate a first parameter based on the independent variable and the mean of the independent variable, encrypt the first parameter, and send the encrypted first parameter to the first data terminal; Receive the encrypted second parameter sent by the first data terminal, where the encrypted second parameter is calculated based on the encrypted first parameter and the difference between the dependent variable and the mean of the dependent variable; Decrypt the encrypted second parameter, and calculate the regression coefficient of the simple linear regression model constructed by the independent variable and the dependent variable according to the decrypted second parameter.
10. The method according to claim 8 or 9, characterized in that Before receiving the difference between the dependent variable and the mean of the dependent variable sent by the first data terminal, it further includes: Perform binning on the independent variable; Receive the sample identifier sent by the first data terminal and the encrypted sample label value corresponding to the sample identifier; Statistically analyze the encrypted sample label values of each bin according to the sample identifier to obtain the encrypted sample label statistical value, and send the encrypted sample label statistical value to the first data terminal for decryption to obtain the dependent variable.
11. A first data terminal server, characterized in that, Applied to a univariate processing system, the univariate processing system includes the first data terminal and the second data terminal, where the first data terminal stores the dependent variable and the first set of feature variables, the second data terminal stores the second set of feature variables, and the second set of feature variables includes the independent variable; the first data terminal includes a first processing module, a first sending module, and a first receiving module: Among them, the first processing module is used to obtain the difference between the dependent variable and the mean of the dependent variable, and send the difference to the second data terminal through the first sending module, and the difference is used to calculate the regression coefficient of the simple linear regression model constructed by the independent variable and the dependent variable; The first receiving module is used to receive the third parameter and the encrypted fourth parameter sent by the second data terminal, where the third parameter is calculated based on the regression coefficient and the mean of the independent variable, and the fourth parameter is calculated based on the regression coefficient and the independent variable; The first processing module is further used to obtain the constant term of the simple linear regression model according to the third parameter and the mean of the dependent variable; obtain the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term, and the dependent variable, and obtain the sum of squared deviations according to the difference between the dependent variable and the mean of the dependent variable; obtain the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and send the encrypted first correlation coefficient to the second data terminal through the first sending module for decryption; The first receiving module is further used to receive the decrypted correlation coefficient sent by the second data terminal and output the correlation coefficient; The first data terminal is further used for: Obtain the calculated value of the second input variable according to the second input variable and the second model parameter, and send the calculated value of the second input variable to the second data terminal; Among them, the second data terminal selects the independent variable in the second feature variable set as the target variable, the other feature variables in the second feature variable set as the first input variable, the first feature variable set as the second input variable, the second data terminal obtains the predicted value of the target variable according to the calculated value of the second input variable, the first input variable and the first model parameter, obtains the sum of squared residuals according to the dependent variable and the predicted value of the target variable, and obtains the sum of squared deviations according to the dependent variable and the mean value of the dependent variable, determines the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and outputs the second correlation coefficient.
12. A second data terminal server, characterized in that, Applied to a univariate processing system, the univariate processing system includes a first data terminal and the second data terminal, wherein the first data terminal stores a dependent variable and a first feature variable set, the second data terminal stores a second feature variable set, and the second feature variable set includes an independent variable; the second data terminal includes a second processing module, a second sending module and a second receiving module; Among them, the second receiving module is used to receive the difference between the dependent variable sent by the first data terminal and the mean value of the dependent variable; The second processing module is used to calculate the regression coefficient of the unary linear regression model constructed by the independent variable and the dependent variable according to the difference; obtain the third parameter according to the regression coefficient and the mean value of the independent variable, obtain the fourth parameter according to the regression coefficient and the independent variable, and perform encryption processing on the fourth parameter; The second sending module is used to send the third parameter and the encrypted fourth parameter to the first data terminal, and the third parameter and the encrypted fourth parameter are used to calculate the encrypted correlation coefficient corresponding to the independent variable; The second receiving module is further used to receive the encrypted first correlation coefficient sent by the first data terminal, decrypt the encrypted first correlation coefficient through the second processing module, and send the decrypted first correlation coefficient to the first data terminal through the second sending module; The second data terminal is further used for: Select the independent variable in the second feature variable set as the target variable, the other feature variables in the second feature variable set as the first input variable, and the first feature variable set as the second input variable; Receive the calculated value of the second input variable sent by the first data terminal, wherein the calculated value of the second input variable is calculated according to the second input variable and the second model parameter; Obtain the predicted value of the target variable according to the calculated value of the second input variable, the first input variable and the first model parameter; Obtain the sum of squared residuals according to the dependent variable and the predicted value of the target variable, and obtain the sum of squared deviations according to the dependent variable and the mean value of the dependent variable; Determine the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and output the second correlation coefficient.
13. A single-variable processing system, characterized in that, Includes a first data terminal and a second data terminal; Wherein, the first data terminal is used to execute the method according to any one of claims 1-7, and the second data terminal is used to execute the method according to any one of claims 8-10.
14. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1-10 when executing the programs stored on the memory.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-10.
16. A variable screening method, characterized in that, Applied to a variable screening system, the variable screening system includes a fifth data terminal and a sixth data terminal. Among them, the fifth data terminal stores a dependent variable and a first independent variable set, and the sixth data terminal stores a second independent variable set; the method includes: The fifth data terminal obtains the first correlation coefficient between each independent variable in the first independent variable set and the dependent variable; The fifth data terminal obtains the difference between the dependent variable and the mean value of the dependent variable, and sends the difference to the sixth data terminal; The sixth data terminal receives the difference, and calculates the regression coefficient of the unary linear regression model constructed by any independent variable in the first independent variable set and the dependent variable according to the difference; The sixth data terminal obtains a third parameter according to the regression coefficient and the mean value of the independent variable, obtains a fourth parameter according to the regression coefficient and the independent variable, performs encryption processing on the fourth parameter, and sends the third parameter and the encrypted fourth parameter to the first data terminal; The fifth data terminal receives the third parameter and the encrypted fourth parameter, obtains the constant term of the unary linear regression model according to the third parameter and the mean value of the dependent variable, obtains the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term, and the dependent variable, and obtains the total sum of squared deviations according to the difference between the dependent variable and the mean value of the dependent variable; obtains the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the total sum of squared deviations, and sends the encrypted first correlation coefficient to the sixth data terminal; The sixth data terminal receives the encrypted first correlation coefficient, decrypts the encrypted first correlation coefficient, and sends the decrypted first correlation coefficient to the fifth data terminal; The fifth data terminal receives the decrypted first correlation coefficient and outputs the first correlation coefficient; iteratively execute the step of obtaining the difference between the dependent variable and the mean value of the dependent variable until the first correlation coefficient corresponding to each independent variable in the second independent variable set is output; Select the independent variables in the first independent variable set whose first correlation coefficients meet the first preset condition to form a first candidate data set, and select the independent variables in the second independent variable set whose first correlation coefficients meet the first preset condition to form a second candidate data set; Select any independent variable in the first candidate data set as the target variable, the other independent variables in the first candidate data set as the first input variables, and the second candidate data set as the second input variables; The sixth data terminal obtains the calculated value of the second input variables according to the second input variables and the second model parameters, and sends the calculated value of the second input variables to the fifth data terminal; The fifth data terminal obtains the predicted value of the target variable based on the calculated value of the second input variable, the first input variable, and the first model parameter; obtains the sum of squared residuals based on the target variable and the predicted value of the target variable, and obtains the total sum of squares based on the target variable and the mean value of the target variable; determines the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the total sum of squares, and outputs the second correlation coefficient; Iteratively execute the step of selecting any independent variable in the first candidate data set as the target variable until the second correlation coefficient corresponding to each independent variable in the first candidate data set when it is used as the target variable is output; Select any independent variable in the second candidate data set as the target variable, other independent variables in the second candidate data set as the third input variable, and the first candidate data set as the fourth input variable; The fifth data terminal obtains the calculated value of the fourth input variable based on the fourth input variable and the fourth model parameter, and sends the calculated value of the fourth input variable to the sixth data terminal; The sixth data terminal obtains the predicted value of the target variable based on the calculated value of the fourth input variable, the third input variable, and the third model parameter; obtains the sum of squared residuals based on the target variable and the predicted value of the target variable, and obtains the total sum of squares based on the target variable and the mean value of the target variable; determines the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the total sum of squares, and outputs the second correlation coefficient; Iteratively execute the step of selecting any independent variable in the second candidate data set as the target variable until the second correlation coefficient corresponding to each independent variable in the second candidate data set when it is used as the target variable is output; Select the independent variables in the first candidate data set whose second correlation coefficients meet the second preset condition to form the third candidate data set, and select the independent variables in the second candidate data set whose second correlation coefficients meet the second preset condition to form the fourth candidate data set. The third candidate data set and the fourth candidate data set form the final candidate data set.
17. A variable screening method, characterized in that, Applied to a variable screening system, the variable screening system includes a fifth data terminal and a sixth data terminal. Among them, the fifth data terminal stores the dependent variable and the first set of independent variables, and the sixth data terminal stores the second set of independent variables; the method includes: Select any independent variable in the first set of independent variables as the target variable, other independent variables in the first set of independent variables as the first input variable, and the second set of independent variables as the second input variable; The sixth data terminal obtains the calculated value of the second input variable based on the second input variable and the second model parameter, and sends the calculated value of the second input variable to the fifth data terminal; The fifth data terminal obtains the predicted value of the target variable based on the calculated value of the second input variable, the first input variable, and the first model parameter; obtains the sum of squared residuals based on the target variable and the predicted value of the target variable, and obtains the total sum of squares based on the target variable and the mean value of the target variable; determines the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the total sum of squares, and outputs the second correlation coefficient; Iteratively execute the step of selecting any independent variable from the first set of independent variables as the target variable until the second correlation coefficients corresponding to each independent variable in the first set of independent variables as the target variable are output; Select any independent variable from the second set of independent variables as the target variable, other independent variables in the second set of independent variables as the third input variables, and the second set of independent variables as the fourth input variables; The fifth data terminal obtains the calculated value of the fourth input variables according to the fourth input variables and the fourth model parameters, and sends the calculated value of the fourth input variables to the sixth data terminal; The sixth data terminal obtains the predicted value of the target variable according to the calculated value of the fourth input variables, the third input variables and the third model parameters; obtains the sum of squared residuals according to the target variable and the predicted value of the target variable, and obtains the sum of squared deviations according to the target variable and the mean value of the target variable; determines the second correlation coefficient corresponding to the target variable according to the sum of squared residuals and the sum of squared deviations, and outputs the second correlation coefficient; Iteratively execute the step of selecting any independent variable from the second set of independent variables as the target variable until the second correlation coefficients corresponding to each independent variable in the second set of independent variables as the target variable are output; Select the independent variables in the first set of independent variables whose second correlation coefficients meet the second preset condition to form the fifth candidate data set, and select the independent variables in the second set of independent variables whose second correlation coefficients meet the second preset condition to form the sixth candidate data set; The fifth data terminal obtains the first correlation coefficients between each independent variable in the fifth candidate data set and the dependent variable; The fifth data terminal obtains the difference between the dependent variable and the mean value of the dependent variable, and sends the difference to the sixth data terminal; The sixth data terminal receives the difference, and calculates the regression coefficient of the simple linear regression model constructed by any independent variable in the sixth candidate data set and the dependent variable according to the difference; The sixth data terminal obtains the third parameter according to the regression coefficient and the mean value of the independent variable, obtains the fourth parameter according to the regression coefficient and the independent variable, encrypts the fourth parameter, and sends the third parameter and the encrypted fourth parameter to the fifth data terminal; The fifth data terminal receives the third parameter and the encrypted fourth parameter, obtains the constant term of the simple linear regression model according to the third parameter and the mean value of the dependent variable, obtains the encrypted sum of squared residuals according to the encrypted fourth parameter, the constant term and the dependent variable, and obtains the sum of squared deviations according to the difference between the dependent variable and the mean value of the dependent variable; obtains the encrypted first correlation coefficient corresponding to the independent variable according to the encrypted sum of squared residuals and the sum of squared deviations, and sends the encrypted first correlation coefficient to the sixth data terminal; The sixth data terminal receives the encrypted first correlation coefficient, decrypts the encrypted first correlation coefficient, and sends the decrypted first correlation coefficient to the fifth data terminal; The fifth data terminal receives the decrypted first correlation coefficient and outputs the first correlation coefficient; Iteratively execute the step of obtaining the difference between the dependent variable and the mean value of the dependent variable until the first correlation coefficients corresponding to each independent variable in the sixth candidate data set are output; Select the independent variables in the fifth candidate data set whose first correlation coefficient meets the first preset condition to form a seventh candidate data set, and select the independent variables in the sixth candidate data set whose first correlation coefficient meets the first preset condition to form an eighth candidate data set. The seventh candidate data set and the eighth candidate data set form the final candidate data set.
18. A variable screening system, characterized in that, It includes a fifth data terminal and a sixth data terminal; wherein, the fifth data terminal and the sixth data terminal are used to execute the method described in claim 16 or 17.
Citation Information
Patent Citations
Characteristic variable analysis method and device, computer equipment and storage medium
CN113934983A