Method and system for predicting signing willingness of rental house, computer device and storage device

By collecting and analyzing the various characteristic data of rental housing users, building and integrating the Logistic regression model, the problem of insufficient accuracy and universality of the prediction of the intention to sign a rental public housing in the existing research is solved, and more efficient prediction and decision support is achieved.

CN119991266APending Publication Date: 2025-05-13SHANGHAI TUOXI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510071147.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

Smart Images

  • Figure CN119991266A_ABST
    Figure CN119991266A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of rental houses, and provides a rental house signing willingness prediction method and system, a computer device and a storage device.The method comprises the steps that basic information, individual characteristics, system factors, housing current situations, housing demand willingness and consumption habits of users are collected to serve as explanatory variables, and signing willingness of the users is collected to serve as dependent variables; the method comprises the following steps: preprocessing user data, carrying out dimension reduction processing by adopting principal component analysis, training a Logistic regression model by utilizing the data subjected to dimension reduction to obtain a regression coefficient, carrying out heterovariance test, carrying out Bootstrap resampling on a sample to train a plurality of Logistic regression models, and fusing model prediction results through a majority voting mechanism to obtain a model prediction result; and analyzing a regression coefficient based on a fused model prediction result, evaluating the influence direction and degree of each explanatory variable on the signing willingness, and determining main factors with significant influence. According to the method, the problems that main factors influencing the contract signing willingness of the rental housing cannot be accurately evaluated and the accuracy of model prediction is insufficient in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rental housing, and in particular to a method, system, computer device and storage device for predicting willingness to sign a rental housing contract. Background Art

[0002] With the acceleration of my country's urbanization process, the demand for housing among urban residents is increasing, especially the demand for rental housing among low- and middle-income groups has increased significantly. However, current research on the rental housing market focuses on the impact of macro factors such as economic factors and social factors, while ignoring the impact of tenants' micro-individual characteristics and emotional factors on their willingness to sign contracts. This leads to certain limitations in existing research when predicting residents' willingness to rent public housing.

[0003] In traditional research, the factors that affect residents' willingness to sign rental housing contracts are usually simplified to external conditions such as rental prices, geographical location, and infrastructure, while ignoring the individual characteristics and emotional preferences of tenants. In fact, tenants' decision-making behavior is complex and dynamic, not only affected by the external environment, but also closely related to internal factors such as their personal emotions, family structure, and career status. It is currently difficult to accurately assess the main factors that affect the willingness to sign rental housing contracts. In addition, tenants' demand and preferences for rental housing will change over time, which makes models based on small-scale data and micro-individual analysis appear incapable of explaining and predicting large-scale rental market behavior, resulting in existing research being limited in the accuracy and universality of prediction models.

[0004] On the other hand, when exploring the important factors that affect residents' willingness to sign rental housing contracts, most existing studies have failed to combine the statistical theoretical basis of large-scale data. Due to the lack of sufficient statistical support, existing models often cannot achieve high-precision predictions during application. This not only limits the effectiveness of rental housing policy formulation, but also affects the scientificity and rationality of decision-making in actual management.

[0005] In general, the current research on the willingness to sign rental housing contracts has the following shortcomings: first, it ignores the diversity of individual characteristics and the influence of emotional factors; second, it lacks statistical theoretical support based on big data, which affects the accuracy and stability of model predictions. Summary of the invention

[0006] Based on this, the purpose of the present invention is to provide a method, system, computer device and storage device for predicting the willingness to sign a rental housing contract, so as to fundamentally solve the existing problems of being unable to accurately evaluate the main factors affecting the willingness to sign a rental housing contract and the insufficient accuracy of model prediction.

[0007] A method for predicting willingness to sign a lease for a house according to an embodiment of the present invention includes:

[0008] The basic information, individual characteristics, institutional factors, housing status, housing demand intention and consumption habits of each user are collected as explanatory variables, and the user's willingness to sign a contract is collected as the dependent variable for binary classification;

[0009] Preprocess the discrete variables and continuous variables in the collected user data, and use principal component analysis to reduce the dimension of the preprocessed data;

[0010] The constructed Logistic regression model is trained using the data after dimensionality reduction, the regression coefficients that affect the willingness to sign a rental housing contract are obtained, and the trained Logistic regression model is tested for heteroscedasticity.

[0011] The collected samples are resampled by Bootstrap using the Bagging method in ensemble learning to generate multiple sample subsets. A Logistic regression model is trained on each sample subset, and the prediction results of multiple Logistic regression models are fused through the majority voting mechanism.

[0012] Based on the prediction results after model fusion, the regression coefficient of the Logistic regression model is analyzed to determine the main explanatory variables that have a significant impact on the willingness to sign a rental housing contract. The direction and degree of the impact of each explanatory variable on the willingness to sign a contract are evaluated based on the positive and negative values ​​and size of the regression coefficient.

[0013] In addition, the method for predicting the willingness to sign a lease for a house according to the above embodiment of the present invention may also have the following additional technical features:

[0014] Furthermore, the step of preprocessing the discrete variables and continuous variables in the collected user data includes:

[0015] Discretize and classify continuous variables according to predetermined intervals, and assign corresponding numerical labels to each interval to convert continuous variables into discrete variables;

[0016] Different categories of discrete variables are numerically encoded according to predetermined rules to organize the discrete variables into numerical format;

[0017] One-Hot encoding is used for the sorted discrete variables, and each category of discrete variables is converted into one-hot encoding.

[0018] Furthermore, the step of performing dimensionality reduction processing on the preprocessed data using principal component analysis includes:

[0019] Calculate the mean of each feature column of the sample data matrix and subtract it from the corresponding feature column data to obtain the centered data matrix;

[0020] Calculate the covariance matrix of the centralized data matrix, and perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues ​​and corresponding eigenvectors of the covariance matrix;

[0021] Sort the eigenvectors in descending order according to the size of the eigenvalues, and select the first k eigenvectors with the largest eigenvalues ​​as the principal component eigenvectors after dimensionality reduction, where k is the preset dimensionality reduction dimension;

[0022] The centered sample data matrix is ​​projected onto the selected first k principal component eigenvectors to obtain the reduced-dimensional data matrix.

[0023] Furthermore, the steps of using the data after dimensionality reduction to train the constructed Logistic regression model, obtaining the regression coefficient that affects the willingness to sign a rental housing contract, and performing a heteroscedasticity test on the trained Logistic regression model include:

[0024] The reduced-dimensional data matrix is ​​input into the Logistic regression model as the independent variable, and the user's willingness to sign a contract is used as the dependent variable to construct a Logistic regression model for predicting the willingness to sign a contract for rental housing.

[0025] The parameters of the Logistic regression model were estimated by the maximum likelihood estimation method to obtain the regression coefficients;

[0026] The obtained regression coefficients are used to evaluate the impact of each explanatory variable on the willingness to sign a rental housing contract, and the positive or negative correlation between the explanatory variable and the willingness to sign a contract and the degree of its impact are determined based on the sign and size of the regression coefficients.

[0027] The heteroscedasticity test was performed on the trained Logistic regression model to evaluate whether the residual term in the Logistic regression model had heteroscedasticity.

[0028] Furthermore, the step of performing heteroscedasticity test on the trained Logistic regression model includes:

[0029] Calculate the residual between the actual value and the predicted probability in the logistic regression model;

[0030] Use Breusch-Pagan test or White test to analyze the variance of residuals to determine whether the variance of residuals changes with the change of independent variables;

[0031] If the result is that heteroskedasticity exists, the weighted least squares method is used to refit the Logistic regression model, and a weight is assigned to each observation, where the size of the weight is inversely proportional to the variance of its residual;

[0032] The regression coefficient and its standard error of the adjusted Logistic regression model were recalculated, and model diagnosis was performed on the adjusted Logistic regression model to verify whether the adjusted Logistic regression model improved the prediction performance and the reliability of the regression coefficient after eliminating heteroskedasticity.

[0033] Furthermore, the steps of using the Bagging method in ensemble learning to perform Bootstrap resampling on the collected samples to generate multiple sample subsets, training a Logistic regression model on each sample subset, and fusing the prediction results of multiple Logistic regression models through a majority voting mechanism include:

[0034] Randomly select samples from the collected sample set by using the Bootstrap resampling method to generate multiple sample subsets, where the sample subsets can be repeatedly sampled;

[0035] Train an independent Logistic regression model on each sample subset, where each Logistic regression model is trained based on the feature data of the corresponding sample subset;

[0036] For each sample to be predicted, multiple Logistic regression models obtained through training are used to predict it and multiple prediction results are obtained;

[0037] The prediction results of multiple Logistic regression models are fused through the majority voting mechanism, and the category that appears most frequently in the prediction results is selected as the final prediction category.

[0038] Furthermore, the steps of analyzing the regression coefficient of the Logistic regression model based on the prediction results after model fusion, determining the main explanatory variables that have a significant impact on the willingness to sign a rental housing contract, and evaluating the direction and degree of influence of each explanatory variable on the willingness to sign a contract based on the positive and negative values ​​and size of the regression coefficient include:

[0039] The prediction results of multiple Logistic regression models obtained after model fusion are combined to obtain the final prediction result;

[0040] Extract the regression coefficient of each explanatory variable from the final Logistic regression model, where the regression coefficient reflects the influence of each explanatory variable on the willingness to sign a rental housing contract;

[0041] Based on the positive and negative values ​​of the regression coefficients, the correlation between each explanatory variable and the willingness to sign the contract is determined;

[0042] Based on the size of the regression coefficient, the influence of each explanatory variable on the willingness to sign the contract is evaluated;

[0043] Combined with the positive and negative values ​​and size of the regression coefficient, the main explanatory variables with significant influence are determined.

[0044] Another embodiment of the present invention aims to provide a rental housing contract signing intention prediction system, the system comprising:

[0045] The data collection module is used to collect each user's basic information, individual characteristics, institutional factors, housing status, housing demand intention and consumption habits as explanatory variables, and collect the user's willingness to sign a contract as the dependent variable for binary classification;

[0046] The data processing module is used to pre-process the discrete variables and continuous variables in the collected user data, and use principal component analysis to perform dimensionality reduction on the pre-processed data;

[0047] The model training module is used to train the constructed Logistic regression model using the data after dimensionality reduction processing, obtain the regression coefficient that affects the willingness to sign a rental housing contract, and perform a heteroscedasticity test on the trained Logistic regression model;

[0048] The model processing module is used to perform Bootstrap resampling on the collected samples using the Bagging method in ensemble learning to generate multiple sample subsets, train a Logistic regression model on each sample subset, and fuse the prediction results of multiple Logistic regression models through a majority voting mechanism;

[0049] The model output module is used to analyze the regression coefficient of the Logistic regression model based on the prediction results after model fusion, determine the main explanatory variables that have a significant impact on the willingness to sign a rental housing contract, and evaluate the direction and degree of influence of each explanatory variable on the willingness to sign a contract based on the positive and negative values ​​and size of the regression coefficient.

[0050] Another embodiment of the present invention aims to provide a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method for predicting willingness to sign a contract for a rental house as described above when executing the computer program.

[0051] Another embodiment of the present invention aims to provide a storage device, characterized in that the storage device stores a computer program, and the computer program can be executed to implement the steps of the method for predicting the willingness to sign a contract for a rental house as described above.

[0052] The method for predicting willingness to sign a rental house contract provided by the embodiment of the present invention pre-processes discrete variables and continuous variables in user data and uses principal component analysis to reduce dimensionality, thereby reducing redundant information in the data, thereby effectively alleviating the dimensionality disaster problem. After dimensionality reduction, the main information of the data can be retained, the model training efficiency is improved, and the overfitting problem caused by too many variables is avoided. By using the Bagging method in ensemble learning to perform Bootstrap resampling on the samples, multiple training subsets are generated, and independent Logistic regression models are trained on each subset. The prediction results of multiple models are fused through a majority voting mechanism, which can effectively improve the stability and accuracy of the model, reduce the risk of overfitting of a single model, and significantly improve the generalization ability of the model on complex data sets. By analyzing the regression coefficient of the Logistic regression model, combined with the fusion of the model The prediction results can clearly evaluate the direction and degree of influence of each explanatory variable on the willingness to sign the contract, and the positive and negative values ​​and size of the regression coefficient can accurately reveal which factors are the key driving factors of the willingness to sign the contract, providing decision makers with accurate analysis basis; by performing heteroscedasticity test on the trained Logistic regression model, detecting and correcting the variance inconsistency problem (heteroscedasticity) of the residual, the accuracy and reliability of the model can be further improved. The model is refitted by weighted least squares method. After eliminating heteroscedasticity, the prediction effect of the model and the stability of the regression coefficient are optimized. At the same time, this method can not only provide accurate prediction of the willingness to sign the contract, but also help understand the influence of different factors on the willingness to sign the contract through the analysis of the regression coefficient. By sorting and evaluating these influencing factors, it can provide a scientific basis for relevant policy formulation, marketing and leasing strategies, etc., and improve the effectiveness and accuracy of decision-making. It solves the existing problems of being unable to accurately evaluate the main factors affecting the willingness to sign the contract for rental housing and the lack of accuracy of model prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flow chart of the method for predicting the willingness to sign a contract for a rental house in the first embodiment of the present invention;

[0054] Figure 2 It is a structural diagram of the rental housing contract signing intention prediction system in the second embodiment of the present invention;

[0055] The following specific implementations will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0056] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0057] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.

[0059] Embodiment 1

[0060] See also Figure 1 , which shows a method for predicting the willingness to sign a contract for a rental house in a first embodiment of the present invention. For the sake of convenience, only the part related to the embodiment of the present invention is shown. The method for predicting the willingness to sign a contract for a rental house provided by the embodiment of the present invention includes:

[0061] Step S10, collecting each user's basic information, individual characteristics, institutional factors, housing status, housing demand intention and consumption habits as explanatory variables, and collecting the user's willingness to sign a contract as a dependent variable for binary classification;

[0062] In one embodiment of the present invention, user data is collected by offline questionnaires, system registered users and tenant data, etc. The offline questionnaire is to collect personal basic information, housing status, consumption habits and other data through questionnaires. System registered user data is obtained from the rental housing system. Tenant data is tenant data with existing lease contracts, which is used to analyze their housing needs and willingness to sign contracts. At this time, the collected user data mainly includes six groups of characteristic variables, including basic information, personality characteristics, institutional factors, housing status, housing demand intention and consumption habits, as well as the user's willingness to sign contracts. In the process of predicting the willingness to sign contracts for public rental housing, the reason for collecting these six groups of characteristic variables is that they fully reflect the key factors affecting tenants' decisions. Specifically, the basic information includes age, gender, marital status, etc., that is, the basic information is the basic socio-demographic characteristics, which can affect tenants' housing needs and choices. Personality characteristics include education level, family size, etc., and personality characteristics may affect tenants' housing demand preferences and living habits. Institutional factors include policies, household registration, etc., and rental housing may be affected by policies. These factors directly affect tenants' right to choose and willingness to sign contracts. The housing status includes the tenant's current housing situation (such as the type of current housing, housing area) and living satisfaction. The housing status directly affects whether the tenant needs to change housing. The housing demand intention includes the tenant's demand type for rental housing, such as the tenant's expected rent, housing location, and area. The housing demand intention reflects the tenant's future housing demand and is an important reference for predicting whether he will sign a contract. The consumption habits include income level, rent affordability, etc., and their consumption habits directly determine whether the tenant can afford the rental housing, which is the economic basis that affects their willingness to sign a contract. In the Logistic regression model, the above six groups of characteristic variables are regarded as independent variables (explanatory variables) to predict the dependent variable (i.e., the user's willingness to sign a contract). Each group of characteristic variables may have an impact on the willingness to sign a contract, so it needs to be fully collected and analyzed. The user's willingness to sign a contract is collected as the dependent variable for binary classification, where the classification is binary (0 is unwilling to sign a contract, 1 is willing to sign a contract). At this time, the above data includes discrete variables (such as gender, marital status, etc.) and continuous variables (such as age, income, housing area, etc.).

[0063] Among them, it should be pointed out that in order to ensure the legality and compliance of data collection, avoid infringing personal privacy or harming public interests, the security and privacy protection of data collection can be achieved in the following various ways in the embodiments of the present invention: Method 1: Data anonymization and de-identification, in which when collecting user data, personal identity information (such as name, ID number, telephone number, etc.) should be anonymized so that the collected data cannot be directly traced back to a specific individual. Specifically, for example, sensitive information directly associated with the user can be removed, and the user information can be encoded into a random unique identifier. Or remove key information that may be used to identify an individual to ensure that even if the data set is leaked, the specific identity of the user cannot be inferred from the data. Method 2: Collect non-sensitive information, focusing on collecting non-sensitive data related to the research target. For example, data such as housing type, income range, age range, marital status, and residential satisfaction are all personal attribute information, but do not involve sensitive privacy. After these data are de-identified, it is not easy to identify the specific individual identity. And avoid collecting highly sensitive data, such as the user's specific address, financial account information, health status, and other information. Method 3: Transparency in data use. Data collection must follow the principles of legality, legitimacy and necessity. The user must be informed in advance of the purpose of data collection, scope of use and storage period, and the user's explicit consent is required. Users have the right to choose whether to provide data and to what extent to share personal information. For fields involving privacy, users must explicitly agree before collection. For example, a privacy statement or agreement is added to the questionnaire or registration system, and the user can continue to provide data only after reading and agreeing. Method 4: Comply with data privacy protection laws and regulations. Data collection and processing should strictly comply with relevant privacy protection laws and regulations. Before data collection, ensure that all applicable privacy and data protection laws and regulations have been followed to ensure legal compliance. Therefore, while doing a good job of anonymizing, de-identifying, and encrypting data storage, strictly follow data privacy protection laws and regulations, and collect non-sensitive and necessary variables. In the embodiment of the present invention, through reasonable technical and management measures, it is possible to ensure the realization of research objectives without hindering the public interest or infringing on personal privacy, so that data collection can avoid infringing on user privacy and achieve compliant data collection and use.

[0064] Step S20, preprocessing the discrete variables and continuous variables in the collected user data, and performing dimensionality reduction processing on the preprocessed data using principal component analysis;

[0065] In one embodiment of the present invention, the step of preprocessing the discrete variables and continuous variables in the collected user data includes:

[0066] Discretize and classify continuous variables according to predetermined intervals, and assign corresponding numerical labels to each interval to convert continuous variables into discrete variables;

[0067] Different categories of discrete variables are numerically encoded according to predetermined rules to organize the discrete variables into numerical format;

[0068] One-Hot encoding is used for the sorted discrete variables, and each category of discrete variables is converted into one-hot encoding.

[0069] Among them, the continuous variable is converted into a discrete variable by performing a "binning" operation on it, that is, dividing it into several intervals, each interval corresponds to a category, and assigning a numerical label to each interval. Specifically, first set one or more intervals according to the distribution of the data and actual needs. For example, the interval of the age variable is set as: [20-30), [30-40), [40-50), etc. Then assign a numerical label to each interval. For example, samples with an age of [20-30) can be assigned to label 1, samples with an age of [30-40) can be assigned to label 2, and so on. At this time, the value of the continuous variable (for example, age) is divided according to the set interval, and the corresponding interval label is assigned to each sample, thereby converting the continuous variable into a discrete variable. For discrete variables, such as categorical variables (gender, city, occupation, etc.), they are converted into numerical format by numerical encoding according to predetermined rules. Among them, a common method is to use "label encoding", that is, assigning an integer value to each category. Specifically, all discrete categories in the data are identified, and then a unique integer value is assigned to each category, and the original category value is replaced with the corresponding numerical label. Furthermore, the sorted discrete variables are encoded using One-Hot encoding to convert the discrete variables into binary vectors. At this time, each category corresponds to an independent binary feature, thereby converting the discrete variables of each category into a one-hot encoding format. For each category, only one feature has a value of 1, and the values ​​of the remaining features are 0.

[0070] Among them, in the process of data processing, discrete variables need to be converted into numerical format so that they can be input into most machine learning algorithms for processing. Among them, since common discrete variables cannot be used directly for calculation, encoding conversion is required, and One-Hot encoding is suitable for discrete variables in feature extraction, and these discrete variables can be converted into data forms suitable for machine learning algorithms, so that the model can be processed correctly. Therefore, in the embodiment of the present invention, by adopting the One-Hot encoding method, not only the dimension of the data is enriched, but also the generalization ability of the model is improved, which lays a solid foundation for subsequent model training and prediction. At the same time, the One-Hot encoding method can also avoid the sorting problem introduced by numerical encoding, and allow the model to better capture the differences between categories.

[0071] Furthermore, many variables (such as personality traits, housing needs, consumption habits, etc.) are involved in data collection, and each variable may contain multiple categories or continuous values. When one-hot encoding is used to process categorical variables, the feature dimension will increase. For example, there are three possible marital statuses (unmarried, married, and divorced), which will be converted into three different variables after one-hot encoding, resulting in an increase in the number of features. The feature set formed by the combination of the above six sets of data is often high-dimensional data. There may be multiple variables in each set of data. Especially when the amount of collected data is large and the variable dimensions are many, high-dimensional data processing requires greater computing resources, and the modeling process will be more complicated. At the same time, high-dimensional data will increase the risk of model overfitting, that is, the model performs well on training data, but has poor generalization ability on new data. In addition, many features may be related, and redundant information increases the complexity of the model, leading to multicollinearity problems between features. Therefore, it is necessary to reduce the high-dimensional features of the data to low-dimensional features, remove redundant information, and retain the main information, thereby improving the computational efficiency and accuracy of the model. Therefore, in the embodiment of the present invention, principal component analysis is used to reduce the dimension of the preprocessed data, wherein the principal component analysis converts high-dimensional data into low-dimensional data through linear transformation, while trying to keep the original information of the data, which reduces the complexity of the model, improves the calculation efficiency, and prevents overfitting while keeping the main information of the data. Therefore, by using principal component analysis for dimensionality reduction, the model can focus on the most important features, reduce unnecessary noise data interference, and thus improve the prediction accuracy.

[0072] Specifically, the step of performing dimensionality reduction processing on the preprocessed data using principal component analysis includes:

[0073] Calculate the mean of each feature column of the sample data matrix and subtract it from the corresponding feature column data to obtain the centered data matrix;

[0074] Calculate the covariance matrix of the centralized data matrix, and perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues ​​and corresponding eigenvectors of the covariance matrix;

[0075] Sort the eigenvectors in descending order according to the size of the eigenvalues, and select the first k eigenvectors with the largest eigenvalues ​​as the principal component eigenvectors after dimensionality reduction, where k is the preset dimensionality reduction dimension;

[0076] The centered sample data matrix is ​​projected onto the selected first k principal component eigenvectors to obtain the reduced-dimensional data matrix.

[0077] Specifically, the sample data integrated by one-hot encoding is first converted into a sample data matrix, and then each feature column in the sample data matrix is ​​averaged, specifically by calculating the mean of each feature column of the sample data matrix and subtracting it from the corresponding feature column data to obtain a centralized data matrix so that the mean of each feature is zero. This step ensures that all features have the same scale to avoid affecting the dimensionality reduction results due to different scales of features. Then the covariance matrix of the centralized data matrix is ​​calculated, where the covariance matrix is ​​used to describe the linear relationship between features. In the dimensionality reduction process, principal component analysis (PCA) relies on the covariance matrix to evaluate the correlation between different features. Then the principal components of the data are determined by calculating the eigenvalues ​​and eigenvectors of the covariance matrix. The eigenvalues ​​represent the variance of the principal components, and the eigenvectors represent the projections of the data in these directions. In order to achieve dimensionality reduction, the largest k eigenvalues ​​and their corresponding eigenvectors are selected. The reason for selecting these principal components is that they contain the largest amount of information (variance) in the data. Then, by projecting the centralized sample data matrix onto the selected k principal components, the reduced-dimensional data matrix is ​​obtained, so that the dimension of the original data is reduced to k, the dimensionality of the data is reduced, and the most important feature information in the data is retained, thereby improving the computational efficiency and accuracy of the model.

[0078] Step S30, using the data after dimensionality reduction processing to train the constructed Logistic regression model, obtain the regression coefficient that affects the willingness to sign a rental housing contract, and perform a heteroscedasticity test on the trained Logistic regression model;

[0079] In one embodiment of the present invention, the above steps include:

[0080] The reduced-dimensional data matrix is ​​input into the Logistic regression model as the independent variable, and the user's willingness to sign a contract is used as the dependent variable to construct a Logistic regression model for predicting the willingness to sign a contract for rental housing.

[0081] The parameters of the Logistic regression model were estimated by the maximum likelihood estimation method to obtain the regression coefficients;

[0082] The obtained regression coefficients are used to evaluate the impact of each explanatory variable on the willingness to sign a rental housing contract, and the positive or negative correlation between the explanatory variable and the willingness to sign a contract and the degree of its impact are determined based on the sign and size of the regression coefficients.

[0083] The heteroscedasticity test was performed on the trained Logistic regression model to evaluate whether the residual term in the Logistic regression model had heteroscedasticity.

[0084] Specifically, the data matrix after dimensionality reduction is used as the independent variable (feature data), and the user's willingness to sign a contract (usually a binary variable, 1 indicates willingness to sign a contract, and 0 indicates no willingness to sign a contract) is used as the dependent variable to construct a Logistic regression model. By using the Logistic regression model for modeling, it is predicted whether the user has the willingness to sign a contract. Among them, the Logistic regression model is suitable for processing binary classification problems. In the prediction of the willingness to sign a contract for a rental house in the embodiment of the invention, the goal is to predict whether the tenant is willing to sign a contract (0 indicates unwillingness, 1 indicates willingness). It is a typical binary classification problem, so it is suitable for the Logistic regression model, and the Logistic regression model can provide the regression coefficient of each characteristic variable, which is convenient for analyzing the influence of each characteristic variable on the willingness to sign a contract. Among them, a positive regression coefficient indicates that the variable is positively correlated with the willingness to sign a contract, that is, as the variable increases, the tenant is more likely to sign a contract. For example, if the tenant's income level coefficient is positive, it means that the higher the income, the greater the probability of signing a rental housing. And a negative regression coefficient indicates a negative correlation, that is, as the variable increases, the probability of the tenant being unwilling to sign a contract increases. For example, if the coefficient of rent affordability is negative, it means that the greater the rent pressure, the lower the willingness to sign a contract. At this time, the regression coefficients of each characteristic variable are of great help to policy making. By analyzing the size of each regression coefficient, policy makers can understand which factors have an important impact on the willingness to sign a contract for rental housing, thereby providing a basis for policy making. At the same time, the Logistic regression model can output the probability value of the event (willingness to sign a contract) based on the input sample to be predicted (that is, the tenant to be predicted), that is, the probability that a tenant to be predicted is willing to sign a contract for rental housing. The Logistic regression model outputs a probability between 0 and 1. At this time, a threshold (such as 0.5) can usually be set to determine whether the willingness to sign a contract is 1 (willing to sign a contract) or 0 (unwilling to sign a contract). Therefore, the Logistic regression model can also effectively obtain the predicted probability of the willingness to sign a contract for the sample to be predicted (that is, the tenant to be predicted).

[0085] Among them, after the Logistic regression model is constructed, the maximum likelihood estimation (MLE) method is used to fit the regression coefficient of the Logistic regression model. The regression coefficient is solved by an iterative algorithm (such as the Newton method or the quasi-Newton method), where the regression coefficient reflects the degree of influence of each feature on the willingness to sign the contract. Specifically, for the above six explanatory variables, each corresponds to a regression coefficient. At this time, according to the regression coefficients obtained by training, the degree of influence of each explanatory variable on the willingness to sign the contract can be analyzed. At this time, the size of the regression coefficient indicates the degree of influence of the feature on the willingness to sign the contract. The larger the absolute value of the regression coefficient, the more significant the influence. Among them, heteroscedasticity refers to the inconsistency of the residual variance in the model, that is, different values ​​of the independent variables may lead to different residual variances. Specifically in statistical modeling, it is assumed that the residual terms are independent and identically distributed, and the variances should be the same. If the residual variance changes with the change of the independent variable, it is called heteroscedasticity. In the Logistic regression model, if there is heteroskedasticity, it will lead to the following problems: 1. The estimation of the regression coefficient is no longer accurate. The variance of the coefficient obtained by the maximum likelihood estimation method for the parameter estimation of the Logistic regression model may be large, and the efficiency of the estimation will decrease. 2. Incorrect significance test. Heteroskedasticity will affect the standard error of the regression coefficient, which may lead to incorrect conclusions when performing significance tests. Therefore, in order to ensure the effectiveness of the Logistic regression model, it is necessary to test the heteroskedasticity of the model to ensure that the estimation of the regression coefficient is valid and reliable. Specifically, the steps for performing heteroskedasticity test on the trained Logistic regression model include:

[0086] Calculate the residual between the actual value and the predicted probability in the logistic regression model;

[0087] Use Breusch-Pagan test or White test to analyze the variance of residuals to determine whether the variance of residuals changes with the change of independent variables;

[0088] If the result is that heteroskedasticity exists, the weighted least squares method is used to refit the Logistic regression model, and a weight is assigned to each observation, where the size of the weight is inversely proportional to the variance of its residual;

[0089] The regression coefficient and its standard error of the adjusted Logistic regression model were recalculated, and model diagnosis was performed on the adjusted Logistic regression model to verify whether the adjusted Logistic regression model improved the prediction performance and the reliability of the regression coefficient after eliminating heteroskedasticity.

[0090] Specifically, the residuals of the Logistic regression model are first calculated, and then the variance of the residuals is analyzed using statistical test methods, such as the Breusch-Pagan test or the White test, to determine whether the variance of the residuals changes with the change of the independent variable. The Breusch-Pagan test tests the relationship between the residual variance and the independent variable. If the residual variance changes with certain characteristics of the independent variable, it means that heteroskedasticity exists. The White test is similar to the Breusch-Pagan test, but makes more general assumptions about the error terms of the model, and can test various forms of heteroskedasticity. At this time, the heteroskedasticity is determined. If the test results show that heteroskedasticity exists, the Logistic regression model needs to be further adjusted. Specifically, the Logistic regression model is fitted with the weighted least squares method (WLS), and the weights in the model are adjusted to eliminate the influence of heteroskedasticity. At this time, a weight is calculated for each observation, and the size of the weight is usually inversely proportional to the variance of the residual. The logistic regression model is then refitted using weighted least squares, that is, the regression coefficients are estimated by minimizing the weighted log-likelihood function, and then the regression coefficients and the corresponding standard errors are recalculated using weighted least squares. The adjusted logistic regression model is then evaluated using model diagnostic tools (such as residual analysis, AIC / BIC, etc.) to ensure that the model's predictive performance is improved after eliminating heteroscedasticity. Cross-validation, accuracy, recall and other indicators are used to verify whether the adjusted model can improve the predictive effect after eliminating heteroscedasticity and ensure the reliability of the regression coefficients.

[0091] Step S40, using the Bagging method in ensemble learning to perform Bootstrap resampling on the collected samples to generate multiple sample subsets, training a Logistic regression model on each sample subset, and fusing the prediction results of multiple Logistic regression models through a majority voting mechanism;

[0092] In one embodiment of the present invention, the above steps include:

[0093] Randomly select samples from the collected sample set by using the Bootstrap resampling method to generate multiple sample subsets, where the sample subsets can be repeatedly sampled;

[0094] Train an independent Logistic regression model on each sample subset, where each Logistic regression model is trained based on the feature data of the corresponding sample subset;

[0095] For each sample to be predicted, multiple Logistic regression models obtained through training are used to predict it and multiple prediction results are obtained;

[0096] The prediction results of multiple Logistic regression models are fused through the majority voting mechanism, and the category that appears most frequently in the prediction results is selected as the final prediction category.

[0097] Specifically, assuming that the original sample set contains n sample data, n samples are randomly selected from it using the Bootstrap resampling method to generate multiple sample subsets. The size of each sample subset is n (that is, the same as the size of the original sample set), but because it is random sampling and repetitions are allowed, some samples may appear repeatedly in a subset, while some samples may be omitted. At this time, B resamplings are performed to generate B sample subsets, each of which contains n data points. These subsets can have overlapping parts because repeated sampling is allowed during sampling. For each sample subset, an independent Logistic regression model is trained using the subset, and the above process is repeated until the corresponding Logistic regression model is trained for all subsets. At this time, each Logistic regression model is trained based on a different sample subset, and the regression coefficients of each Logistic regression model may be different. Then, for each sample to be predicted, it is input into each trained Logistic regression model. Each Logistic regression model will give a prediction result. According to the prediction probability of each Logistic regression model, it is converted into a binary category label (0 or 1). The prediction results of all Logistic regression models are fused through the majority voting mechanism, and the final prediction label is the category that appears most times in the category label.

[0098] Among them, Bagging (Bootstrap Aggregating) is an integrated learning technology. It generates multiple different data sets by resampling the training data, and trains a model (weak model) on each data set. Finally, the prediction results of each model are fused by majority voting to form a strong model. This method has several important functions in prediction: 1. Reduce variance and improve stability: Since the data set for each training is different, Bagging can effectively reduce the variance of the model and prevent overfitting. Among them, a single model often has a strong dependence on the training data, but by integrating the results of multiple models, a more stable and generalized model can be obtained. 2. Model fusion improves accuracy: Each Logistic regression model may have a certain deviation or variance. By integrating multiple models and averaging, the prediction accuracy can be improved without increasing complexity. 3. Reduce the risk of overfitting of the model: Bagging resamples the samples so that each model is trained on different sub-samples, reducing the model's dependence on specific samples and reducing the risk of overfitting.

[0099] Therefore, by sampling with replacement from the original training data set to generate multiple different training sets, due to the sampling with replacement, there will be overlaps between different sample sets. The bagging method can effectively reduce the impact of some extreme samples on the model. At the same time, since an independent Logistic regression model is trained for each resampled training set, each resampled data set may contain different sample combinations, so that each trained model may be slightly different. At the same time, since there may be certain errors in the Logistic regression model trained in each round of sampling, but these errors are different on different data sets, at this time, by integrating the results of multiple models, the bagging method can integrate the predictions of multiple models, reduce the random errors of a single model, and improve the generalization ability of the overall model. And by using the majority voting mechanism for the models trained on different data sets, the bagging method can smooth the prediction results, reduce the volatility in the prediction, and make the model more robust. Therefore, by fusing the prediction results of each model, a stronger integrated model is finally formed. Bagging can greatly improve the stability and generalization ability of the model through the fusion method of the majority voting mechanism.

[0100] Step S50, based on the prediction results after model fusion, the regression coefficient of the Logistic regression model is analyzed to determine the main explanatory variables that have a significant impact on the willingness to sign a contract for rental housing, and the direction and degree of influence of each explanatory variable on the willingness to sign a contract are evaluated based on the positive and negative values ​​and size of the regression coefficient;

[0101] In one embodiment of the present invention, the above steps include:

[0102] The prediction results of multiple Logistic regression models obtained after model fusion are combined to obtain the final prediction result;

[0103] Extract the regression coefficient of each explanatory variable from the final Logistic regression model, where the regression coefficient reflects the influence of each explanatory variable on the willingness to sign a rental housing contract;

[0104] Based on the positive and negative values ​​of the regression coefficients, the correlation between each explanatory variable and the willingness to sign the contract is determined;

[0105] Based on the size of the regression coefficient, the influence of each explanatory variable on the willingness to sign the contract is evaluated;

[0106] Combined with the positive and negative values ​​and size of the regression coefficient, the main explanatory variables with significant influence are determined.

[0107] Specifically, assume that multiple Logistic regression models have been trained using the Bagging method, and each model will output a prediction result. First, all these prediction results are combined to generate the final prediction result. For each sample to be predicted, each Logistic regression model will give a prediction probability and convert it into a binary classification label (0 or 1) according to a preset threshold (such as 0.5). At this time, the final prediction label is obtained through a majority voting mechanism, that is, the category with the most occurrences in the prediction results of all models is taken as the final prediction category. Then, the regression coefficient of each explanatory variable is extracted from the final Logistic regression model after the model fusion. These regression coefficients reflect the degree of influence of each explanatory variable on the willingness to sign a contract for rental housing. At this time, the correlation is judged by analyzing the positive and negative values ​​of the regression coefficients. If the regression coefficient is greater than 0, it means that there is a positive correlation between the explanatory variable and the willingness to sign a contract, that is, the increase of the variable tends to increase the willingness to sign a contract. If the regression coefficient is less than 0, it means that there is a negative correlation between the explanatory variable and the willingness to sign a contract, that is, the increase of the variable tends to reduce the willingness to sign a contract. At this time, by analyzing the sign (positive or negative) of each regression coefficient, the direction of the influence of each explanatory variable on the willingness to sign a contract can be determined. Then, based on the size of the regression coefficient, the degree of influence is evaluated. At this time, the larger the absolute value of the regression coefficient, the greater the influence of the explanatory variable on the willingness to sign the contract. This is because the size of the regression coefficient is directly related to the magnitude of the change in the target variable (willingness to sign the contract) by the variable. At this time, for each explanatory variable, the size of its regression coefficient is evaluated to determine which variables have the most significant impact on the willingness to sign the contract for rental housing among all the explanatory variables. By analyzing the sign and size of the regression coefficient, combined with statistical significance tests (such as Wald test, z test, etc.), the main explanatory variables with significant impact on the willingness to sign the contract are determined. According to the positive and negative values ​​and size of the regression coefficient, the variables that have a strong impact on the willingness to sign the contract are further identified. These variables are the main explanatory variables of the model and are usually used preferentially for decision analysis or policy making. Through the above analysis, a result list containing the main explanatory variables, their regression coefficients, correlations and degrees of influence is obtained. This list can help understand the importance of different variables in predicting the willingness to sign the contract for rental housing. According to the size of the regression coefficient, the explanatory variables with the greatest impact are ranked first to provide a basis for relevant decisions. For example, if the regression coefficient of a variable is large and positive, it means that the variable may be a key factor in increasing the willingness to sign a contract. Policymakers can intervene in this factor and adjust the rental housing policy in a targeted manner.

[0108] Therefore, by analyzing the regression coefficient of the Logistic regression model, combined with the positive and negative values ​​and size of the regression coefficient, we can determine which explanatory variables have a significant impact on the willingness to sign a rental housing contract. This process not only helps to understand the impact direction of each variable, but also provides a scientific basis for the formulation of relevant policies or decisions, enhances the practical application value of the model, enables the model to more accurately predict the willingness of tenants to sign contracts, and helps the formulation and implementation of rental housing policies.

[0109] In summary, the method for predicting willingness to sign a rental house in the above embodiment of the present invention reduces redundant information in the data by preprocessing discrete variables and continuous variables in user data and reducing the dimension by principal component analysis, thereby effectively alleviating the dimensionality curse problem. After the dimensionality reduction process, the main information of the data can be retained, the model training efficiency is improved, and the overfitting problem caused by too many variables is avoided. By adopting the Bagging method in ensemble learning to perform Bootstrap resampling on the samples, multiple training subsets are generated, and an independent Logistic regression model is trained on each subset. The prediction results of multiple models are fused through the majority voting mechanism, which can effectively improve the stability and accuracy of the model, reduce the risk of overfitting of a single model, and significantly improve the generalization ability of the model on complex data sets. By analyzing the regression coefficient of the Logistic regression model, combined with the model fusion The combined prediction results can clearly evaluate the direction and degree of influence of each explanatory variable on the willingness to sign a contract, and the positive and negative values ​​and size of the regression coefficient can accurately reveal which factors are the key driving factors of the willingness to sign a contract, providing decision makers with accurate analysis basis; by testing the heteroscedasticity of the trained Logistic regression model, detecting and correcting the variance inconsistency problem (heteroscedasticity) of the residual, the accuracy and reliability of the model can be further improved. The model is refitted by weighted least squares method. After eliminating heteroscedasticity, the prediction effect of the model and the stability of the regression coefficient are optimized. At the same time, this method can not only provide accurate predictions of the willingness to sign a contract, but also help understand the influence of different factors on the willingness to sign a contract through the analysis of the regression coefficient. By sorting and evaluating these influencing factors, it can provide a scientific basis for relevant policy formulation, marketing and leasing strategies, and improve the effectiveness and accuracy of decision-making. It solves the existing problems of being unable to accurately evaluate the main factors affecting the willingness to sign a contract for rental housing and the lack of accuracy of model prediction.

[0110] Embodiment 2

[0111] See also Figure 2 , is a structural diagram of a rental housing contract signing intention prediction system provided by the second embodiment of the present invention. For the sake of convenience, only the parts related to the embodiment of the present invention are shown. The rental housing contract signing intention prediction system includes:

[0112] The data collection module 11 is used to collect basic information, individual characteristics, institutional factors, housing status, housing demand intention and consumption habits of each user as explanatory variables, and collect the user's willingness to sign a contract as a dependent variable for binary classification;

[0113] The data processing module 12 is used to pre-process the discrete variables and continuous variables in the collected user data, and perform dimensionality reduction processing on the pre-processed data using principal component analysis;

[0114] The model training module 13 is used to train the constructed Logistic regression model using the data after the dimension reduction process, obtain the regression coefficient that affects the willingness to sign a rental housing contract, and perform a heteroscedasticity test on the trained Logistic regression model;

[0115] The model processing module 14 is used to perform Bootstrap resampling on the collected samples using the Bagging method in ensemble learning to generate multiple sample subsets, train a Logistic regression model on each sample subset, and fuse the prediction results of multiple Logistic regression models through a majority voting mechanism;

[0116] The model output module 15 is used to analyze the regression coefficient of the Logistic regression model based on the prediction results after model fusion, determine the main explanatory variables that have a significant impact on the willingness to sign a rental housing contract, and evaluate the direction and degree of influence of each explanatory variable on the willingness to sign a contract based on the positive and negative values ​​and size of the regression coefficient.

[0117] Furthermore, in one embodiment of the present invention, the data processing module 12 includes:

[0118] A discretization processing unit, used to discretize and classify the continuous variable according to a predetermined interval, and assign a corresponding numerical label to each interval to convert the continuous variable into a discrete variable;

[0119] A numerical coding unit, used for numerically coding discrete variables of different categories according to predetermined rules, so as to organize the discrete variables into numerical formats;

[0120] The encoding processing unit is used to use One-Hot encoding on the sorted discrete variables and convert the discrete variables of each category into one-hot encoding.

[0121] Furthermore, in one embodiment of the present invention, the data processing module 12 includes:

[0122] A centralization processing unit is used to calculate the mean of each feature column of the sample data matrix and subtract it from the corresponding feature column data to obtain a centralized data matrix;

[0123] A covariance processing unit, used to calculate the covariance matrix of the centralized data matrix, and perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues ​​and corresponding eigenvectors of the covariance matrix;

[0124] An eigenvalue sorting and selection unit is used to sort the eigenvectors in descending order according to the size of the eigenvalues, and select the first k eigenvectors with the largest eigenvalues ​​as the principal component eigenvectors after dimensionality reduction, where k is a preset dimensionality reduction dimension;

[0125] The dimensionality reduction processing unit is used to project the centralized sample data matrix onto the selected first k principal component eigenvectors to obtain a data matrix after dimensionality reduction.

[0126] Furthermore, in one embodiment of the present invention, the model training module 13 includes:

[0127] The Logistic regression model building module is used to input the dimension-reduced data matrix as an independent variable into the Logistic regression model, and use the user's signing willingness as the dependent variable to build a Logistic regression model for predicting the signing willingness of rental housing;

[0128] The regression coefficient determination unit is used to estimate the parameters of the Logistic regression model by the maximum likelihood estimation method to obtain the regression coefficient;

[0129] The regression coefficient evaluation unit is used to evaluate the impact of each explanatory variable on the willingness to sign a rental housing contract using the obtained regression coefficient, and to determine the positive or negative correlation between the explanatory variable and the willingness to sign a contract and the degree of its impact based on the sign and size of the regression coefficient;

[0130] The heteroscedasticity test unit is used to perform heteroscedasticity test on the trained Logistic regression model to evaluate whether the residual term in the Logistic regression model has heteroscedasticity.

[0131] Furthermore, in one embodiment of the present invention, the heteroscedasticity test unit is used to:

[0132] Calculate the residual between the actual value and the predicted probability in the logistic regression model;

[0133] Use Breusch-Pagan test or White test to analyze the variance of residuals to determine whether the variance of residuals changes with the change of independent variables;

[0134] If the result is that heteroskedasticity exists, the weighted least squares method is used to refit the Logistic regression model, and a weight is assigned to each observation, where the size of the weight is inversely proportional to the variance of its residual;

[0135] The regression coefficient and its standard error of the adjusted Logistic regression model were recalculated, and model diagnosis was performed on the adjusted Logistic regression model to verify whether the adjusted Logistic regression model improved the prediction performance and the reliability of the regression coefficient after eliminating heteroskedasticity.

[0136] Furthermore, in one embodiment of the present invention, the model processing module 14 includes:

[0137] A sample subset selection unit is used to randomly select samples from the collected sample set by using a Bootstrap resampling method to generate multiple sample subsets, wherein the sample subsets can be repeatedly sampled;

[0138] A model training unit, used for training an independent Logistic regression model on each sample subset, wherein each Logistic regression model is trained based on the feature data of the corresponding sample subset;

[0139] The model prediction unit is used to predict each sample to be predicted using multiple Logistic regression models obtained through training to obtain multiple prediction results;

[0140] The model fusion unit is used to fuse the prediction results of multiple Logistic regression models through a majority voting mechanism and select the category with the most occurrences in the prediction results as the final prediction category.

[0141] Furthermore, in one embodiment of the present invention, the model output module 15 includes:

[0142] A prediction result acquisition unit is used to integrate the prediction results of multiple Logistic regression models obtained after model fusion to obtain the final prediction result;

[0143] The regression coefficient extraction unit is used to extract the regression coefficient of each explanatory variable from the final Logistic regression model, where the regression coefficient reflects the influence of each explanatory variable on the willingness to sign a rental housing contract;

[0144] A correlation judgment unit is used to judge the correlation between each explanatory variable and the willingness to sign a contract based on the positive and negative values ​​of the regression coefficient;

[0145] The signing willingness evaluation unit is used to evaluate the influence of each explanatory variable on the signing willingness based on the size of the regression coefficient;

[0146] The explanatory variable determination unit is used to determine the main explanatory variables with significant influence by combining the positive and negative values ​​and size of the regression coefficient.

[0147] The implementation principle and technical effects of the rental housing contract signing intention prediction system provided in the embodiment of the present invention are the same as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding contents in the aforementioned method embodiment.

[0148] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method for predicting the willingness to sign a contract for a rental house as described in the first embodiment above are implemented.

[0149] An embodiment of the present invention further provides a storage device, wherein the storage device stores a computer program, and the computer program can be executed to implement the steps of the method for predicting the willingness to sign a contract for a rental house as described in the first embodiment above.

[0150] Exemplarily, the computer program can be divided into one or more modules, one or more modules are stored in the memory and executed by the processor to complete the present invention. One or more modules can be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program in the computer device. For example, the computer program can be divided into the steps of the rental housing contract signing intention prediction method provided by each of the above method embodiments.

[0151] Those skilled in the art will appreciate that the above description of the computer device is merely an example and does not constitute a limitation on the computer device. The computer device may include more or fewer components than described above, or a combination of certain components, or different components, such as input and output devices, network access devices, buses, etc.

[0152] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device, and uses various interfaces and lines to connect various parts of the entire computer device.

[0153] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the computer device by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as an interface display function, an interface interaction function, etc.), etc.; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0154] If the module / unit integrated in the computer device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above method embodiment, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above method embodiment when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, electrical signal and software distribution medium, etc.

[0155] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0156] The above-described embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.

Claims

1. A method for predicting willingness to sign a rental house contract, characterized in that: The method comprises: The basic information, individual characteristics, institutional factors, housing status, housing demand intention and consumption habits of each user are collected as explanatory variables, and the user's willingness to sign a contract is collected as the dependent variable for binary classification; Preprocess the discrete variables and continuous variables in the collected user data, and use principal component analysis to reduce the dimension of the preprocessed data; The constructed Logistic regression model is trained using the data after dimensionality reduction, the regression coefficients that affect the willingness to sign a rental housing contract are obtained, and the trained Logistic regression model is tested for heteroscedasticity. The collected samples are resampled by Bootstrap using the Bagging method in ensemble learning to generate multiple sample subsets. A Logistic regression model is trained on each sample subset, and the prediction results of multiple Logistic regression models are fused through the majority voting mechanism. Based on the prediction results after model fusion, the regression coefficient of the Logistic regression model is analyzed to determine the main explanatory variables that have a significant impact on the willingness to sign a rental housing contract. The direction and degree of the impact of each explanatory variable on the willingness to sign a contract are evaluated based on the positive and negative values ​​and size of the regression coefficient.

2. The method for predicting willingness to sign a lease for a rental house according to claim 1, characterized in that: The step of preprocessing the discrete variables and continuous variables in the collected user data includes: Discretize and classify continuous variables according to predetermined intervals, and assign corresponding numerical labels to each interval to convert continuous variables into discrete variables; Different categories of discrete variables are numerically encoded according to predetermined rules to organize the discrete variables into numerical format; One-Hot encoding is used for the sorted discrete variables, and each category of discrete variables is converted into one-hot encoding.

3. The method for predicting willingness to sign a lease for a rental house according to claim 1, characterized in that: The step of performing dimensionality reduction processing on the preprocessed data by using principal component analysis comprises: Calculate the mean of each feature column of the sample data matrix and subtract it from the corresponding feature column data to obtain the centered data matrix; Calculate the covariance matrix of the centralized data matrix, and perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues ​​and corresponding eigenvectors of the covariance matrix; Sort the eigenvectors in descending order according to the size of the eigenvalues, and select the first k eigenvectors with the largest eigenvalues ​​as the principal component eigenvectors after dimensionality reduction, where k is the preset dimensionality reduction dimension; The centered sample data matrix is ​​projected onto the selected first k principal component eigenvectors to obtain the reduced-dimensional data matrix.

4. The method for predicting willingness to sign a lease for a rental house according to claim 1, characterized in that: The steps of training the constructed Logistic regression model using the data after dimensionality reduction processing, obtaining the regression coefficient affecting the willingness to sign a rental housing contract, and performing a heteroscedasticity test on the trained Logistic regression model include: The reduced-dimensional data matrix is ​​input into the Logistic regression model as the independent variable, and the user's willingness to sign a contract is used as the dependent variable to construct a Logistic regression model for predicting the willingness to sign a contract for rental housing. The parameters of the Logistic regression model were estimated by the maximum likelihood estimation method to obtain the regression coefficients; The obtained regression coefficients are used to evaluate the impact of each explanatory variable on the willingness to sign a rental housing contract, and the positive or negative correlation between the explanatory variable and the willingness to sign a contract and the degree of its impact are determined based on the sign and size of the regression coefficients. The heteroscedasticity test was performed on the trained Logistic regression model to evaluate whether the residual term in the Logistic regression model had heteroscedasticity.

5. The method for predicting willingness to sign a lease for a rental house according to claim 4, characterized in that: The step of performing heteroscedasticity test on the trained Logistic regression model includes: Calculate the residual between the actual value and the predicted probability in the logistic regression model; Use Breusch-Pagan test or White test to analyze the variance of residuals to determine whether the variance of residuals changes with the change of independent variables; If the result is that heteroskedasticity exists, the weighted least squares method is used to refit the Logistic regression model, and a weight is assigned to each observation, where the size of the weight is inversely proportional to the variance of its residual; The regression coefficient and its standard error of the adjusted Logistic regression model were recalculated, and model diagnosis was performed on the adjusted Logistic regression model to verify whether the adjusted Logistic regression model improved the prediction performance and the reliability of the regression coefficient after eliminating heteroskedasticity.

6. The method for predicting willingness to sign a lease for a rental house according to claim 1, characterized in that: The steps of using the Bagging method in ensemble learning to perform Bootstrap resampling on the collected samples to generate multiple sample subsets, training a Logistic regression model on each sample subset, and fusing the prediction results of multiple Logistic regression models through a majority voting mechanism include: Randomly select samples from the collected sample set by using the Bootstrap resampling method to generate multiple sample subsets, where the sample subsets can be repeatedly sampled; Train an independent Logistic regression model on each sample subset, where each Logistic regression model is trained based on the feature data of the corresponding sample subset; For each sample to be predicted, multiple Logistic regression models obtained through training are used to predict it and multiple prediction results are obtained; The prediction results of multiple Logistic regression models are fused through the majority voting mechanism, and the category that appears most frequently in the prediction results is selected as the final prediction category.

7. The method for predicting willingness to sign a lease for a rental house according to claim 1, characterized in that: The steps of analyzing the regression coefficient of the Logistic regression model based on the prediction results after model fusion, determining the main explanatory variables that have a significant impact on the willingness to sign a contract for rental housing, and evaluating the direction and degree of influence of each explanatory variable on the willingness to sign a contract based on the positive and negative values ​​and size of the regression coefficient include: The prediction results of multiple Logistic regression models obtained after model fusion are combined to obtain the final prediction result; Extract the regression coefficient of each explanatory variable from the final Logistic regression model, where the regression coefficient reflects the influence of each explanatory variable on the willingness to sign a rental housing contract; Based on the positive and negative values ​​of the regression coefficients, the correlation between each explanatory variable and the willingness to sign the contract is determined; Based on the size of the regression coefficient, the influence of each explanatory variable on the willingness to sign the contract is evaluated; Combined with the positive and negative values ​​and size of the regression coefficient, the main explanatory variables with significant influence are determined.

8. A rental housing contract signing intention prediction system, characterized in that: The system comprises: The data collection module is used to collect each user's basic information, individual characteristics, institutional factors, housing status, housing demand intention and consumption habits as explanatory variables, and collect the user's willingness to sign a contract as the dependent variable for binary classification; The data processing module is used to pre-process the discrete variables and continuous variables in the collected user data, and use principal component analysis to perform dimensionality reduction on the pre-processed data; The model training module is used to train the constructed Logistic regression model using the data after dimensionality reduction processing, obtain the regression coefficient that affects the willingness to sign a rental housing contract, and perform a heteroscedasticity test on the trained Logistic regression model; The model processing module is used to perform Bootstrap resampling on the collected samples using the Bagging method in ensemble learning to generate multiple sample subsets, train a Logistic regression model on each sample subset, and fuse the prediction results of multiple Logistic regression models through a majority voting mechanism; The model output module is used to analyze the regression coefficient of the Logistic regression model based on the prediction results after model fusion, determine the main explanatory variables that have a significant impact on the willingness to sign a rental housing contract, and evaluate the direction and degree of influence of each explanatory variable on the willingness to sign a contract based on the positive and negative values ​​and size of the regression coefficient.

9. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for predicting the willingness to sign a contract for a rental house as claimed in any one of claims 1 to 7 are implemented.

10. A storage device, characterized in that: The storage device stores a computer program, which can be executed to implement the steps of the method for predicting the willingness to sign a contract for a rental house as described in any one of claims 1 to 7.