A Value Model for Bankruptcy Disposal

Through the bankruptcy disposal value model, using tax data synchronization and linear regression model, it solves the problems of time-consuming and error-prone data collection in traditional property valuation methods, and realizes efficient, accurate and real-time property valuation, which is suitable for the field of real estate valuation.

CN118780823BActive Publication Date: 2025-09-30珠海华发金融科技研究院有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410799745.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2025-09-30
Estimated Expiration
2044-06-20

AI Technical Summary

Technical Problem

Traditional property valuation methods have the problems of time-consuming data collection, prone to errors, non-real-time, high cost, difficulty in processing unstructured data, and lack of automation and intelligence, which lead to inaccurate valuations.

Method used

A bankruptcy disposal value model is adopted, including data collection module, data cleaning and preprocessing module, data segmentation module, model training module, model evaluation module, model optimization module, deployment and use module and monitoring and update module. Tax data synchronization and linear regression model are used to estimate the value of real estate, combined with caching technology and permission protection to achieve automated and real-time valuation.

Benefits of technology

It improves the efficiency and accuracy of data collection, reduces human errors, provides highly interpretable valuation results, has high computational efficiency and good stability, and can continuously optimize the model and adapt to data changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118780823B_ABST
    Figure CN118780823B_ABST
Patent Text Reader

Abstract

This invention discloses a bankruptcy disposal value model, comprising a data collection module, a data cleaning and preprocessing module, a data segmentation module, a model training module, a model evaluation module, a model optimization module, a deployment and utilization module, and a monitoring and update module. The data collection module includes a tax data synchronization unit and a historical data observation unit. The tax data synchronization unit synchronizes property taxes and fees daily through a big data interface with the tax authorities and infers the property value based on these taxes and fees. The historical data observation unit stores observable data in a time series and provides historical data query capabilities. This invention uses tax data as a benchmark, resulting in significantly higher accuracy. It uses a linear regression model for testing, making it more scientific. Continuous improvement and optimization result in a more suitable model. It also includes detailed logging and error notifications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bankruptcy disposal, and in particular to a value model for bankruptcy disposal. Background Art

[0002] Property valuation is a crucial component of bankruptcy resolution. Property valuation is the process of determining the market value of real estate (such as real estate). Currently, there are multiple methods available for valuation. Traditional valuation methods have advantages and disadvantages, including inaccurate estimates and ambiguous reference bases. Linear regression based on deep learning can effectively address these issues. The following is a comparison of some common property valuation methods:

[0003] 1. Market Comparison Approach:

[0004] This is a common property valuation method, particularly applicable to real estate; it determines market value by comparing the subject property with similar properties that have been sold or leased; the key is selecting the right comparable properties and then adjusting their prices or rents to account for any differences between the subject property and the comparable properties.

[0005] advantage:

[0006] Intuitive: Easy to understand because it is based on transaction data already available in the market.

[0007] Practicality: Applicable to a variety of real estate types, especially residential and commercial properties.

[0008] Market reaction: reflects the current market supply and demand relationship and investor sentiment.

[0009] shortcoming:

[0010] Data limitations: Sales or rental data on a large number of comparable properties is required, and sometimes it is difficult to obtain sufficient comparable data.

[0011] Adjustments for differences: Adjustments for differences between comparable properties and the target property can be subjective and affect the accuracy of the valuation.

[0012] Data quality: This depends on the quality and accuracy of the data; errors in the data may lead to inaccurate valuations.

[0013] 2. Income Approach:

[0014] The income approach, applicable to rental properties, estimates the property's value based on the rental income or future cash flows it generates; this approach includes sub-methods such as direct capitalization, discounted cash flow, and net operating income.

[0015] advantage:

[0016] Suitable for rental properties: For rental properties, especially commercial properties, accurate valuations can be provided.

[0017] Consider future cash flow: This takes into account the future rental income or cash flow of the property, making it more suitable for investment considerations.

[0018] shortcoming:

[0019] High data requirements: Accurate rent and expense data are required, as well as reasonable capitalization rates or discount rates.

[0020] Assumption Sensitivity: It is very sensitive to assumptions such as rental growth rate and discount rate. Inaccurate assumptions may lead to inaccurate valuations.

[0021] Not applicable to owner-occupied properties: Not applicable to owner-occupied properties and some special-purpose properties.

[0022] 3. Cost Approach

[0023] The cost approach is used to estimate a property's value. It's based on the property's rebuilding cost or replacement cost, then adjusted to account for depreciation and market conditions. This approach is often used for special-use properties, such as industrial facilities and historic buildings.

[0024] advantage:

[0025] For special purpose properties: For special purpose properties or new construction properties, useful valuation methods are provided.

[0026] Actual cost based: Based on reconstruction or replacement cost, which takes into account the actual construction cost of the property.

[0027] shortcoming:

[0028] Complex depreciation adjustments: Accurate adjustments for property depreciation need to be made, which may involve subjective judgment.

[0029] Assumption Sensitivity: Very sensitive to assumptions about land depreciation and market conditions.

[0030] Unreflective of market conditions: May not necessarily reflect current market supply and demand and investor sentiment.

[0031] Traditional data collection methods have some shortcomings in some cases, which can be partially addressed in the digital age; the following are some common shortcomings of traditional data collection methods:

[0032] 1) Manual and time-consuming: Traditional data collection usually requires manual collection, organization, and recording of data, which is a time-consuming task. Manual data collection may require a lot of time and human resources.

[0033] 2) Error-prone: Manual data collection is prone to errors and data inaccuracies, especially in the case of large-scale data sets; human factors may lead to data entry errors, omissions or mistakes.

[0034] 3) Not real-time: Traditional data collection may take some time to complete, so the data is not real-time enough; this may cause problems in application areas that require timely decision-making (such as finance, healthcare, etc.).

[0035] 4) High Cost: Manual data collection typically requires hiring and training personnel, purchasing and maintaining equipment, and incurring costs associated with data collection; this can result in high expenses.

[0036] 5) Limited data volume: Manual data collection is usually limited by manpower and time, and therefore cannot process large-scale or high-frequency data; this limits the richness and diversity of the data.

[0037] 6) Difficulty in processing unstructured data: Traditional methods generally have difficulty effectively processing unstructured data such as text, images, and audio; these data types require more complex techniques to extract useful information.

[0038] 7) Difficulty tracking changes: Once a data collection method is established, it is often difficult to easily modify it to accommodate changing requirements or data sources; this can lead to an inflexible data collection process.

[0039] 8) Lack of automation and intelligence: Traditional methods lack automation and intelligence capabilities and cannot automatically adapt to patterns or trends in data; this makes data analysis and mining more challenging. Summary of the Invention

[0040] The purpose of the present invention is to provide a value model for bankruptcy disposal to solve the problems raised in the above background technology.

[0041] To achieve the above-mentioned object, the present invention provides the following technical solutions: a value model for bankruptcy disposal, comprising a data collection module, a data cleaning and preprocessing module, a data segmentation module, a model training module, a model evaluation module, a model optimization module, a deployment and use module, and a monitoring and updating module;

[0042] The data collection module includes a tax data synchronization unit and a historical data observation unit. The tax data synchronization unit is used to synchronize property taxes and fees daily through a big data interface with the tax authorities and infer the property value based on the taxes and fees. The historical data observation unit is used to store observable data in time series and provide historical data query.

[0043] The data cleaning and preprocessing module is used to clean and preprocess the data, including processing missing values, outliers and duplicate values. The data cleaning and preprocessing module can also perform feature engineering, extract new features or transform existing features to improve model performance;

[0044] The data segmentation module is used to divide the data set into a training set and a test set, using a ratio of 80% training and 20% testing to facilitate the verification of the model performance during the training process;

[0045] The model training module is used to train the selected machine learning model using the training data, so that the model learns patterns and relationships in the data in order to make valuation predictions;

[0046] The model evaluation module is used to evaluate the results and make visual adjustments by predicting the problems of the data;

[0047] The model optimization module is used to adjust and optimize the model according to the evaluation results to improve the accuracy of the valuation;

[0048] The deployment module is used to deploy the trained and optimized model into a production environment for property valuation. The interface inputs the property's feature data, and the model generates an estimate of the market value.

[0049] The monitoring and updating module is used to regularly update the model based on performance monitoring and data changes. The monitoring and updating module includes retraining the model, adjusting hyperparameters, adding new features, or deleting irrelevant features.

[0050] Preferably, the tax data synchronization unit performs the following synchronization steps:

[0051] a1. First, import the necessary packages;

[0052] a2. Read API data;

[0053] a3. Define data labels;

[0054] a4. Historical data observation: Store observable data in time series and provide historical data query.

[0055] Preferably, the data segmentation module divides the data set into a training set and a test set, wherein the training set is used to determine the parameters of the model and the test set is used to judge the effect of the model.

[0056] Preferably, the model training module uses a linear regression model to train data and predict test set data.

[0057] Preferably, the model optimization module adjusts and optimizes the model according to the evaluation results, including hyperparameter adjustment, feature selection and model integration. The optimization steps of the model optimization module are as follows:

[0058] a) First, try to rebuild the model with the three most correlated features, compare them with the original model, and find the three most correlated features, which are used as independent variables to build the model;

[0059] b) Compare the scores of the predicted test set data to see whether the linear regression model 1 is high or low, and then look at its performance on the entire dataset.

[0060] Preferably, the reconstructing the model includes using multiple algorithms to respectively establish models and compare the models.

[0061] A value model for bankruptcy disposal includes the following process steps:

[0062] Step S1, data interface: synchronize property taxes and fees every day through the big data interface with the tax authorities;

[0063] Step S2: Valuation request: E-Chain or other property platforms initiate a valuation request and pass in property information;

[0064] Step S3: Using cache technology, first query whether there is corresponding cache data in the cache. If not, directly call the model for valuation;

[0065] Step S4: Calculate the valuation result based on linear regression;

[0066] Step S5: Finally, an independent cache layer is generated to cache the data;

[0067] Step S6: Record the request log and the generated log;

[0068] Step S7: To ensure that there are no security vulnerabilities from XSS malicious code attacks, strict permission protection is required for the request. If necessary, before execution, its content must be analyzed and identified to prevent malicious code from being included. If any malicious code is included, the request will not be executed and a log will be recorded.

[0069] Step S8: Save the estimated data and model in time series and continue to improve them.

[0070] A value model for bankruptcy disposal includes a computing program execution status monitoring module, which is used to provide a UI interface to monitor the execution status of all computing program methods in real time, including: status, time of last successful execution, and abnormal error information.

[0071] Compared with the prior art, the present invention has the following beneficial effects:

[0072] 1) Provide benchmark data or forecast data for observation by relevant parties, such as the E-chain system, general real estate, and real estate agents.

[0073] 2) Improve the performance of collecting data and reduce human errors.

[0074] 3) Strong interpretability: The parameters of the linear regression model are intuitively interpretable. The coefficient of each feature represents the degree of influence of the feature on the property, which helps to explain the results of the model and provide insights to stakeholders.

[0075] 4) High computational efficiency: The training and prediction process of linear regression is very fast, especially when the number of features is relatively small and the amount of data is not large.

[0076] 5) Good stability: Linear regression is relatively stable to noise and outliers in the data because it estimates parameters through the least squares method.

[0077] 6) Continuously optimize the model through long-term data observation and calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 It is a schematic diagram of the process of the present invention;

[0079] Figure 2 This is a schematic diagram of the E-chain flow of the present invention. DETAILED DESCRIPTION

[0080] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0081] See also Figure 1-2 ,The present invention provides a technical solution: a value model for bankruptcy disposal, including a data collection module, a data cleaning and preprocessing module, a data segmentation module, a model training module, a model evaluation module, a model optimization module, a deployment and use module, and a monitoring and update module;

[0082] The data collection module includes a tax data synchronization unit and a historical data observation unit. The tax data synchronization unit is used to synchronize property taxes and fees daily through a big data interface with the tax authorities and infer the property value based on the taxes and fees. The historical data observation unit is used to store observable data in time series and provide historical data query.

[0083] The data cleaning and preprocessing module is used to clean and preprocess the data, including processing missing values, outliers and duplicate values. The data cleaning and preprocessing module can also perform feature engineering, extract new features or transform existing features to improve model performance;

[0084] def load_data():

[0085] #Read the file separated by spaces and turn it into a continuous array

[0086] firstdata=np.fromfile('housing.data',sep=")

[0087] #Add attributes

[0088] feature_names=['CRIM','ZN','INDUS','CHAS','NOX','RM','AGE','DIS','RAD','TAX','PTRATIO','B','LSTAT','MEDV']

[0089] # Column length

[0090] feature_num=len(feature_names)

[0091] #print(firstdata.shape) output: (7084,)

[0092] #print(firstdata.shape[0] / / feature_nums) Output: 506

[0093] #Construct a 506*14 two-dimensional array

[0094] data=firstdata.reshape([firstdata.shape[0] / / feature_num,feature_num])

[0095] The data segmentation module is used to divide the data set into a training set and a test set, using a ratio of 80% training and 20% testing to facilitate the verification of the model performance during the training process;

[0096] The model training module is used to train the selected machine learning model using the training data, so that the model learns patterns and relationships in the data in order to make valuation predictions;

[0097] The model evaluation module is used to evaluate the results and make visual adjustments based on the problems with the predicted data. For example, if the predicted score is around 76% and the mean square error (RMSE) is around 4.5, visual adjustments are made to better identify the problems with the predicted data.

[0098] df_coef = pd.DataFrame()

[0099] df_coef['Title']=data.columns.delete(-1)

[0100] df_coef['Coef'] = coef

[0101] df_coef

[0102] Evaluation Model:

[0103] plt.scatter(y_test,line_pre,label='y')

[0104] plt.plot([y_test.min(),y_test.max()],[y_test.min(),y_test.max()],'k--',lw=4,label='predicted')

[0105] The model optimization module is used to adjust and optimize the model according to the evaluation results to improve the accuracy of the valuation;

[0106] The deployment module is used to deploy the trained and optimized model into a production environment for property valuation. The interface inputs the property's feature data, and the model generates an estimate of the market value.

[0107] The monitoring and updating module is used to regularly update the model based on performance monitoring and data changes. The monitoring and updating module includes retraining the model, adjusting hyperparameters, adding new features, or deleting irrelevant features.

[0108] In the present invention, the tax data synchronization unit performs the following synchronization steps:

[0109] a1. First import the necessary packages:

[0110] import pandas as pd

[0111] import numpy as np

[0112] import matplotlib.pyplot as plt

[0113] import seaborn as sns

[0114] from sklearn.model_selection import train_test_split

[0115] from sklearn.linear_model import LinearRegression

[0116] from sklearn.metrics import mean_squared_error

[0117] plt.style.use('ggplot')

[0118] %load_ext klab-autotime;

[0119] a2. Read API data;

[0120] data=pd.read_data('.. / data_files / sync_data_housing / data.csv')

[0121] data.info()

[0122] a3. Define data labels;

[0123]

[0124]

[0125] a4. Historical data observation: Store observable data in time series and provide historical data query.

[0126] In the present invention, the data segmentation module divides the data set into a training set and a test set, wherein the training set is used to determine the parameters of the model, and the test set is used to judge the effect of the model.

[0127] #The training set is set to 80% of the total data

[0128] ratio=0.8

[0129] offset=int(data.shape[0]*ratio)

[0130] training_data=data[:offset]

[0131] #print(training_data.shape)

[0132] #axis=0 means column

[0133] #axis=1 means row

[0134] #\ indicates a line break, no need to enter

[0135] maximums,minimums,avgs=training_data.max(axis=0),training_data.min(axis=0),training_data.sum(axis=0) / \training_data.shape[0]

[0136] #View the maximum, minimum, and average values ​​of each column in the training set

[0137] #print(maximums,minimums,avgs)

[0138] # Normalize all data

[0139] for iin range(feature_num):

[0140] #print(maximums[i],minimums[i],avgs[i])

[0141] #Normalization, subtracting the mean is to remove the common part and highlight individual differences

[0142] data[:,i]=(data[:,i]-avgs[i]) / (maximums[i]-minimums[i])

[0143] # Cover the above training set

[0144] training_data=data[:offset]

[0145] #The remaining 20% ​​is the test set

[0146] test_data = data[offset:]

[0147] return training_data,test_data

[0148] In the present invention, the model training module uses a linear regression model to train data and predict test set data.

[0149] linear_model=LinearRegression()

[0150] linear_model.fit(X_train,y_train)

[0151] coef=linear_model.coef_#regression coefficient

[0152] line_pre=linear_model.predict(X_test)

[0153] print('SCORE:{:.4f}'.format(linear_model.score(X_test,y_test)))

[0154] print('RMSE:{:.4f}'.format(np.sqrt(mean_squared_error(y_tes t,line_pre))))

[0155] coef

[0156] In the present invention, the model optimization module adjusts and optimizes the model according to the evaluation results, including hyperparameter adjustment, feature selection and model integration. The optimization steps of the model optimization module are as follows:

[0157] a) First try to reconstruct the model using the three most correlated features and compare it with the original model;

[0158] df.corr()['MEDV'].abs().sort_values(ascending=False).head(4)

[0159] The three most correlated features were obtained and used as independent variables to build the model;

[0160] b) Compare the scores of the predicted test set data to see whether the linear regression model 1 is high or low, and then look at the performance of the entire dataset;

[0161] X2=np.array(data[['LSTAT','RM','PIRATIO']])

[0162] X2_train,X2_test,y_train,y_test=train_test_split(X2,y,random_state=1,test_size=0.2)

[0163] linear_model2=LinearRegression()

[0164] linear_model2.fit(X2_train,y_train)

[0165] print(linear_model2.intercept_)

[0166] print(linear_model2.coef_)

[0167] line2_pre=linear_model2.predict(X2_test)#predicted value

[0168] print('SCORE:{:.4f}'.format(linear_model2.score(X2_test,y_test)))#Model score

[0169] print('RMSE:{:.4f}'.format(np.sqrt(mean_squared_error(y_test,line2_pre))))#RMSE (standard error)

[0170] In the present invention, the model reconstruction includes using multiple algorithms to respectively establish models and compare the models.

[0171] GradientBoosting

[0172] from sklearn import ensemble

[0173] #params={'n_estimators':500,'max_depth':4,'min_samples_split':1,'learning_rate':0.01,'loss':'ls'}

[0174] #clf=ensemble.GradientBoostingRegressor(**params)

[0175] clf=ensemble.GradientBoostingRegressor()

[0176] clf.fit(X_train,y_train)

[0177] clf_pre=clf.predict(X_test)#predicted value

[0178] print('SCORE:{:.4f}'.format(clf.score(X_test,y_test)))#Model score

[0179] print('RMSE:{:.4f}'.format(np.sqrt(mean_squared_error(y_test,clf_pre))))#RMSE (standard error)

[0180] Return of Lasso (Least Absolute Shrinkage and Selection Operator)

[0181] from sklearn.linear_model import Lasso

[0182] lasso=Lasso()

[0183] lasso.fit(X_train,y_train)

[0184] y_predict_lasso=lasso.predict(X_test)

[0185] r2_score_lasso=r2(y_test,y_predict_lasso)

[0186] print('SCORE:{:.4f}'.format(lasso.score(X_test,y_test)))#Model score

[0187] print('RMSE:{:.4f}'.format(np.sqrt(mean_squared_error(y_test,y_predict_lasso))))#RMSE (standard error)

[0188] print('The R-squared value of the Lasso model is:',r2_score_lasso)

[0189] ElasticNet Regression

[0190] enet=ElasticNet()

[0191] enet.fit(X_train,y_train)

[0192] y_predict_enet = enet.predict(X_test)

[0193] r2_score_enet = r2(y_test, y_predict_enet)

[0194] print('SCORE:{:.4f}'.format(enet.score(X_test, y_test))) # Model score

[0195] print('RMSE:{:.4f}'.format(np.sqrt(mean_squared_error(y_tes t, y_predict_enet)))) # RMSE (Standard error)

[0196] print("The R-squared value of the ElasticNet model is:", r2_score_enet)

[0197] Support Vector Regression (SVR)

[0198] from sklearn.linear_model import ElasticNet

[0199] from sklearn.svm import SVR

[0200] from sklearn.metrics import confusion_matrix, classification_report

[0201] from sklearn.metrics import r2_score as r2, mean_squared_error as mse, mean_absolute_error as mae

[0202] def svr_model(kernel):

[0203] svr = SVR(kernel = kernel)

[0204] svr.fit(X_train, y_train)

[0205] y_predict = svr.predict(X_test)

[0206] #score():Returns the coefficient of determination R^2of theprediction.

[0207] print(kernel,'The default evaluation value of SVR is:',svr.score(X_test,y_test))

[0208] print(kernel,'The R-squared value of SVR is:',r2(y_test,y_predict))

[0209] print(kernel,'The mean squared error of SVR is:',mse(y_test,y_predict))

[0210] print(kernel,'The mean absolute error of SVR is:',mae(y_test,y_predict))

[0211] #print(kernel,'The mean squared error of SVR is:',mse(scalery.inverse_transform(y_test),scalery.inverse_transform(y_predict)))

[0212] #print(kernel,'The mean absolute error of SVR is:',mae(scalery.inverse_transform(y_test),scalery.inverse_transform(y_predict)))

[0213] In the data analysis process, feature design is the most important. The quality of data analysis results actually depends mainly on the features. This is also the case in the project, which is constantly being optimized and fine-tuned.

[0214] A value model for bankruptcy disposal includes the following process steps:

[0215] Step S1, data interface: synchronize property taxes and fees every day through the big data interface with the tax authorities;

[0216] Step S2: Valuation request: E-Chain or other property platforms initiate a valuation request and pass in property information;

[0217] Step S3: Using cache technology, first query whether there is corresponding cache data in the cache. If not, directly call the model for valuation;

[0218] Step S4: Calculate the valuation result based on linear regression;

[0219] line2_pre_all = linreg2.predict(X2) #predicted value

[0220] print('SCORE:{:.4f}'.format(linreg2.score(X2,y)))#Model score

[0221] print('RMSE:{:.4f}'.format(np.sqrt(mean_squared_error(y,line2_pre_all))))#RMSE (standard error)

[0222] Step S5: Finally, an independent cache layer is generated to cache the data;

[0223] Step S6: Record the request log and the generated log;

[0224] Step S7: To ensure that there are no security vulnerabilities from XSS malicious code attacks, strict permission protection is required for the request. If necessary, before execution, its content must be analyzed and identified to prevent malicious code from being included. If any malicious code is included, the request will not be executed and a log will be recorded.

[0225] Step S8: Save the estimated data and model in time series and continue to improve them.

[0226] A value model for bankruptcy disposal includes a computing program execution status monitoring module, which is used to provide a UI interface to monitor the execution status of all computing program methods in real time, including: status, time of last successful execution, and abnormal error information.

[0227] The present invention: synchronizes property taxes and fees every day through a big data interface with the tax authorities; initiates a valuation request from E-Link or other property platforms and transmits property information; adopts caching technology to first query whether there is corresponding cached data in the cache, and if not, directly calls the model for valuation; calculates the valuation result based on linear regression; finally generates an independent cache layer to cache the data; records the request log and the generated log; in order to ensure that there are no security vulnerabilities such as XSS malicious code attacks, strict permission protection is required for the request, and if necessary, its content is analyzed and identified before execution, and malicious code is not allowed to be included. If malicious code is included, it will not be executed and a log will be recorded; the estimated data and model are saved in time series and continuously improved.

[0228] The valuation model of the present invention is based on the latest Python and tax benchmark data, making full use of today's most advanced Tensorflow computing technology and mature and stable caching technology (such as: Redis real-time database and MQTT message queue), and has great advantages in the fields of big data, cloud platforms, real estate networks, etc.

[0229] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in this field. Although the embodiments of the present invention have been shown and described, it is understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A bankruptcy disposal value model, characterized by: It includes data collection module, data cleaning and preprocessing module, data segmentation module, model training module, model evaluation module, model optimization module, deployment and use module, and monitoring and update module; The data collection module includes a tax data synchronization unit and a historical data observation unit. The tax data synchronization unit is used to synchronize property taxes and fees daily through a big data interface with the tax authorities and infer the property value based on the taxes and fees. The historical data observation unit is used to store observable tax data in time series and provide historical data query. The data cleaning and preprocessing module is used to clean and preprocess the data, including processing missing values, outliers and duplicate values. The data cleaning and preprocessing module extracts new features or transforms existing features to improve model performance; The data segmentation module is used to divide the data set into a training set and a test set, using a ratio of 80% training and 20% testing to facilitate the verification of the model performance during the training process; The model training module is used to train the selected machine learning model using the training data, so that the model learns patterns and relationships in the data in order to make valuation predictions; The model evaluation module is used to evaluate the results and make visual adjustments by predicting the problems of the data; The model optimization module is used to adjust and optimize the model according to the evaluation results to improve the accuracy of the valuation; The deployment module is used to deploy the trained and optimized model into a production environment for property valuation. The interface inputs the property's feature data, and the model generates an estimate of the market value. The monitoring and updating module is used to regularly update the model based on performance monitoring and data changes. The monitoring and updating module includes retraining the model, adjusting hyperparameters, adding new features, or deleting irrelevant features.

2. The bankruptcy disposal value model according to claim 1, characterized in that: The synchronization steps of the tax data synchronization unit are as follows: a1. First, import the necessary packages; a2. Read API data; a3. Define data labels; a4. Historical data observation: Store observable data in time series and provide historical data query.

3. The bankruptcy disposal value model according to claim 1, characterized in that: The data segmentation module divides the data set into a training set and a test set, wherein the training set is used to determine the parameters of the model, and the test set is used to evaluate the effect of the model.

4. The bankruptcy disposal value model according to claim 1, characterized in that: The model training module uses a linear regression model to train the data and predict the test set data.

5. The bankruptcy disposal value model according to claim 1, characterized in that: The model optimization module adjusts and optimizes the model according to the evaluation results, including hyperparameter adjustment, feature selection and model integration. The optimization steps of the model optimization module are as follows: a) First, try to rebuild the model with the three most correlated features, compare them with the original model, and find the three most correlated features, which are used as independent variables to build the model; b) Compare the scores of the predicted test set data to see whether the linear regression model 1 is high or low, and then look at its performance on the entire dataset.

6. The bankruptcy disposal value model according to claim 5, characterized in that: The model reconstruction includes using multiple algorithms to respectively establish models and comparing the models.

7. The value model of bankruptcy disposal according to claim 1, characterized in that: The process steps include: Step S1, data interface: synchronize property taxes and fees every day through the big data interface with the tax authorities; Step S2: Valuation request: E-Chain or other property platforms initiate a valuation request and pass in property information; Step S3: Using cache technology, first query whether there is corresponding cache data in the cache. If not, directly call the model for valuation; Step S4: Calculate the valuation result based on linear regression; Step S5: Finally, an independent cache layer is generated to cache the data; Step S6: Record the request log and the generated log; Step S7: To ensure that there are no security vulnerabilities from XSS malicious code attacks, strict permission protection is required for the request. Before execution, its content must be analyzed and identified to prevent malicious code from being included. If any malicious code is included, the request will not be executed and a log will be recorded. Step S8: Save the estimated data and model in time series and continue to improve them.

8. The bankruptcy disposal value model according to claim 1, characterized in that: It includes a computing program execution status monitoring module, which is used to provide a UI interface to monitor the execution status of all computing program methods in real time, including: status, time of last successful execution, and abnormal error information.

Citation Information

Patent Citations

  • Real estate big data-based automatic real estate price assessment system and method

    CN108921597A

  • A real estate evaluation method based on machine learning

    CN109146278A