Land price prediction method and system for small sample city data expansion
By constructing a multi-source index system for land prices and using data from similar cities for sample expansion, the problem of insufficient and uneven sample size in land price prediction in small sample cities is solved, and the prediction accuracy and model adaptability are improved.
Patent Information
- Application Number
- CN202510308233.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
The existing technology faces small sample problems in land price prediction, insufficient sample size and uneven sample, resulting in the model being unable to effectively capture potential patterns in the data, affecting the accuracy and stability of the prediction.
By constructing a multi-source index system for land prices, similar cities in the research area are screened, and similar cities and research areas are used as two independent databases to input land price regression prediction models for model training, and data expansion and model retraining is performed based on associated similar samples.
It effectively solves the problems of insufficient sample size and uneven data, and improves the adaptability of the model and the accuracy of land price prediction.
Smart Images

Figure CN120146308A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of urban land price prediction and sample data expansion, and particularly to a land price prediction method and system for small-sample urban data expansion. Background Art
[0002] Land price prediction plays a crucial role in urban planning, real estate development and government decision-making; accurate land price prediction helps the government formulate reasonable land use policies and optimize resource allocation. In addition, it is also an important tool for real estate development, which can provide scientific investment reference and reduce investment and operation risks. At the same time, land price prediction can also help evaluate the development potential of the city and provide a basis for formulating economic development strategies; through accurate land price prediction, it is possible to better understand market demand and trends, promote the healthy development of the land market, and thus promote the sustainable development of the entire city and region.
[0003] However, current land price prediction methods face many challenges in practical applications. The particularly critical and prominent problem is the small-sample problem. The land transaction sample data of a certain city is scarce, and machine learning prediction cannot be effectively carried out. In addition, it mainly manifests in two aspects: insufficient sample size and sample imbalance. When conducting land price prediction, due to the inactivity of the land transaction market in some cities, it is difficult to obtain complete and comprehensive land transaction data, resulting in a very prominent small-sample problem of insufficient sample size. Moreover, the insufficient sample size is difficult to support the normal operation of machine learning algorithms, resulting in the model being unable to effectively capture the potential patterns in the data, affecting the accuracy and stability of the prediction. Sample imbalance refers to the skewed data distribution, with too few observations for some targets, resulting in poor prediction effects of the model for these minority value range regions. The problem of data imbalance generally exists in land price prediction. This data imbalance poses a huge challenge to deep learning and machine learning algorithms. It often shows better prediction effects for the majority value range regions, while poorer prediction effects for the minority value range regions, and it is difficult to meet the prediction requirements in a small-sample environment.
[0004] The current research direction is to generate virtual data associated with sampling points through model simulation as sample data. When constructing a model in this way, there will be a problem of excessive homogenization of model samples. The specific indicators in the model are the geographical environment factors of the sampling point samples. Generating based on the sampling point samples through model simulation will lead to excessive homogenization of sample data. The virtual sample data cannot fully represent the real data, and even the generated samples may deviate from the actual data characteristics, ultimately resulting in overfitting of the prediction model and low model prediction accuracy. Currently, research also considers using data smoothing techniques to smooth the data distribution to reduce noise and fluctuations. This method has obvious effects when dealing with time series data or data with noise. However, in the case of insufficient sample size, smoothing processing may lead to over-simplification of the data and loss of important feature information. In addition, data smoothing is extremely limited in dealing with complex multi-dimensional data. Land price prediction has particularity. The samples in cities are generally not many, and even the samples in some cities are small. This limits the application of land price prediction, and it is also impossible to construct multi-source indicators that conform to the actual situation (with small samples, no matter how many indicators there are, they are necessarily meaningless), which also limits the completeness of model construction. Currently, evaluations are often made through manual research, and the subjectivity of the evaluation results is relatively large. There is no feasible technical solution idea at present. Summary of the Invention
[0005] The purpose of the present invention is to solve the technical problems pointed out in the background technology, and provide a land price prediction method and system for data expansion of small-sample cities, construct a multi-source index system for land prices that conforms to actual applications, obtain land transaction sample data of the same-level cities in the study area, screen similar cities in the study area based on the multi-source index system for land prices, use the similar cities and the study area as two independent databases to input into the land price regression prediction model for model training of land transaction price prediction respectively, sort and correspond according to the SHAP values of the characteristics of the independent database of the similar cities and the independent database of the study area to screen and construct associated similar samples, expand the data of the study area based on the associated similar samples and retrain the model after expansion, effectively solve the technical problems of insufficient sample size and data imbalance, and improve the model adaptability and land price prediction accuracy.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] A land price prediction method for data expansion of small-sample cities, the method includes:
[0008] S1. Taking the small-sample city as the study area, obtaining the land transaction sample data of the study area and the same-level cities in the study area and performing data preprocessing, and the land transaction sample data includes land transaction prices;
[0009] S2. Construct a multi-source index system for land prices, collect and preprocess the multi-source index data of land prices in the study area and cities of the same level as the study area. The multi-source index system for land prices includes four categories of indicators: geographical location, community living facilities, plot attributes, and regional economic conditions. Each category of indicator includes several indicators; calculate the average value of all indicators in the multi-source index system for land prices in the study area and cities of the same level as the study area, and screen the top N1 cities of the same level as the study area with the smallest Euclidean distance from the study area as the similar cities of the study area;
[0010] S3. Construct a land price regression prediction model with the land transaction price of the land transaction sample data as the dependent variable and all indicators in the multi-source index system for land prices as the independent variables. Input the similar cities and the study area as two independent databases into the land price regression prediction model respectively for separate model training of land transaction price prediction; obtain the feature SHAP values of the two independent database samples in the land price regression prediction model through the random forest algorithm, screen the top N2 identical and / or similar ones in the feature SHAP value ranking from the independent database of the similar cities as the associated similar samples, and based on the associated similar samples, use the independent database of the similar cities to expand the corresponding data of the independent database of the study area and re-train the model after the expansion;
[0011] S4. Collect the multi-source index data of the land price in the study area, input it into the land price regression prediction model after model training with data expansion, and output the predicted land transaction price.
[0012] To better implement the present invention, the land transaction sample data in method S1 includes the multi-source index data corresponding to the multi-source index system for land prices. The collection of the multi-source index data of land prices in method S2 includes associated collection from the land transaction sample data and collection of the remaining data that the land transaction sample data does not include the multi-source index data of land prices. The preprocessing of the multi-source index data of land prices includes data verification processing; the land transaction price of the land transaction sample data is the land transaction unit price.
[0013] Preferably, the next level of the geographical location category index includes the shortest distance from the land sample to the hospital, the shortest distance to the hotel, the shortest distance to the factory, the shortest distance to the square, the shortest distance to the villa area, the shortest distance to the commercial service, the shortest distance to the tourist attraction, and / or the shortest distance to the kindergarten; the next level of the community living facilities category index includes the number of commercial services, the number of kindergartens, the number of hotels, the number of primary schools, the number of medical points, the number of bus stops, the number of convenience stores, the number of bars, and / or the number of shopping centers around the land transaction sample; the next level of the land parcel attribute category index includes population density, normalized difference vegetation index NDVI, minimum floor area ratio, maximum floor area ratio, and / or benchmark land price. The normalized difference vegetation index NDVI is calculated by collecting satellite remote sensing data and extracting the reflectance value NIR in the red light band and the reflectance value Red in the near-infrared band according to the following formula: The regional economic status includes the per capita living consumption expenditure of residents.
[0014] Preferably, the multi-source index data of land prices in similar cities and the study area are respectively projected to the 2000 National Geodetic Coordinate System using Gauss-Kruger projection through ArcMap software and rasterized. Based on the concentration degree of the multi-source index data of land prices, grid cells are clustered and divided, and the position attributes of the grids are correspondingly assigned. The position attribute is the position information of the grid in the grid cell and the similar city or the study area.
[0015] Preferably, in method S2, the multi-source index data of land prices in the study area and cities at the same level as the study area are the data within N3 kilometers around the sample geographical location of the land transaction sample data.
[0016] Preferably, the calculation expression of the Euclidean distance between the study area and cities at the same level as the study area is as follows:
[0017] where d is the Euclidean distance, Y 1 and Y 2 are the land transaction prices of the study area and cities at the same level as the study area respectively, and X 1n and X 2b are the average values of the corresponding index n of the study area and cities at the same level as the study area respectively.
[0018] Preferably, the preprocessing of the multi-source index data of land prices in method S2 includes normalization processing, and the normalization processing expression is as follows: X ij is the data corresponding to index j before normalization processing; i represents the city number of the study area and cities at the same level as the study area; X' ij is the data corresponding to index j after normalization processing; max(X) is the maximum value of the data before normalization processing; min(X) is the minimum value of the data before normalization processing.
[0019] Preferably, the feature SHAP value is obtained through the random forest algorithm model of the interpretable model SHAP value. The random forest model consists of a random forest using no less than 500 decision trees, and the maximum depth of each decision tree is 15.
[0020] Preferably, the screening method for associated similar samples in method S3 is as follows:
[0021] Sort the feature SHAP values of the independent database in the study area and the independent database of similar cities respectively, compare the top P1 feature SHAP values in the sorting, and select the first P2 of the independent database in the study area and the independent database of similar cities that are exactly the same, and the remaining P1 - P2 that are shared by the study area and similar cities as the associated similar samples.
[0022] A land price prediction system for small-sample urban data expansion, including a data acquisition and processing module, a similar city database, a research area database, a multi-source index system for land prices, a similar city screening module, a land price regression prediction model, and an output module. The multi-source index system for land prices includes four category indexes: geographical location, community living facilities, plot attributes, and regional economic conditions. Each category index includes several indexes. The data acquisition and processing module obtains the land transaction sample data and multi-source index data of land prices in the research area, preprocesses them, and stores them in the research area database. The data acquisition and processing module obtains the land transaction sample data and multi-source index data of land prices in the same-level cities of the research area, preprocesses them, and stores them in the database of the same-level cities of the research area. The land transaction sample data includes land transaction prices, and the research area is a small-sample city. The similar city screening module calculates the average value of all indexes in the multi-source index system of land prices in the research area and the same-level cities of the research area, and screens the top N1 same-level cities of the research area with the smallest Euclidean distance from the research area as the similar cities of the research area and stores them in the similar city database. The similar city screening module is constructed with the land transaction price of the land transaction sample data as the dependent variable and all indexes in the multi-source index system of land prices as the independent variables. The similar cities and the research area are used as two independent databases and are respectively input into the land price regression prediction model for separate model training of land transaction price prediction. The feature SHAP values of the two independent database samples in the land price regression prediction model are obtained through the random forest algorithm model. The top N2 identical and / or similar ones in the feature SHAP value ranking of the independent database of the similar cities and the independent database of the research area are selected from the independent database of the similar cities as the associated similar samples. Based on the associated similar samples, the independent database of the research area is expanded with corresponding data and the model is retrained after the expansion. The multi-source index data of land prices in the research area is collected and input into the land price regression prediction model after the model training of the expansion, and the predicted land transaction price is output. The output module is used to output the predicted land transaction price of the research area.
[0023] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0024] (1) The present invention constructs a multi-source index system for land prices that conforms to practical applications, obtains the land transaction sample data of the same-level cities in the research area, screens the similar cities in the research area based on the multi-source index system of land prices, uses the similar cities and the research area as two independent databases and respectively inputs them into the land price regression prediction model for model training of land transaction price prediction, screens and constructs associated similar samples according to the feature SHAP value ranking of the independent database of the similar cities and the independent database of the research area, expands the data of the research area based on the associated similar samples and retrains the model after the expansion, effectively solving the technical problems of insufficient sample size and data imbalance, and improving the model adaptability and land price prediction accuracy.
[0025] (2) Compared with the sample expansion techniques such as traditional data augmentation, transfer learning, oversampling, weight adjustment, and data smoothing, the present invention conducts sample expansion by associating similar samples for the particularity of land price prediction. The associated similar samples are constructed by associating similar attributes such as similar cities, similar land transaction unit prices, the same or similar SHAP values of features, and the same geographical attributes, which have objective attributes conforming to land transactions and are beneficial to improving the prediction accuracy and wide adaptability of the model. The multi-source index system for land price construction in the present invention has an index system containing rich and objectively practical multi-source index data, which is conducive to the deep learning regression processing of the land price regression prediction model and improves the prediction accuracy of the model. Description of the Drawings
[0026] Figure 1 It is the method flow chart of the land price prediction method of the present invention;
[0027] Figure 2 It is the principle of the source data acquisition method for the multi-source index data of land price in the embodiment;
[0028] Figure 3 It is the method flow principle diagram of the random forest algorithm model of the interpretable model SHAP value in the embodiment;
[0029] Figure 4 It is the front and back part sorting schematic diagram of the feature SHAP value of a certain city for example;
[0030] Figure 5 It is the order of the top 10 feature SHAP values of the feature SHAP value of a certain city for example;
[0031] Figure 6 It is the comparison schematic diagram of the prediction result and the real result of using the original data of Zhoushan City to input the unexpanded training land price regression prediction model in the embodiment;
[0032] Figure 7 It is the comparison schematic diagram of the prediction result and the real result of using the expanded data of Zhoushan City to input the land price regression prediction model after expanded training in the embodiment;
[0033] Figure 8 It is the comparison schematic diagram of the prediction result and the real result of using the original data of Ningbo City to input the unexpanded training land price regression prediction model in the embodiment;
[0034] Figure 9 It is the comparison schematic diagram of the prediction result and the real result of using the expanded data of Ningbo City to input the land price regression prediction model after expanded training in the embodiment. Detailed Embodiments
[0035] The present invention will be further described in detail below in conjunction with embodiments:
[0036] Embodiment
[0037] As Figure 1 shown, a land price prediction method for small-sample urban data expansion, the method includes:
[0038] S1. Taking a small-sample city as the research area, obtaining the land transaction sample data of the research area and cities at the same level as the research area and performing data preprocessing. The land transaction sample data includes land transaction prices. The present invention takes a small-sample city as the research area, and the small-sample city shows two situations: insufficient data volume and sufficient data volume but data imbalance. The present invention takes cities at the same level as the research area as the sample source of the research area, gradually screens similar cities and similar samples, balances the small-sample city data samples of the small-sample city, improves the model adaptability of the small-sample city (the sample is too small to apply the model for prediction) and the land price prediction accuracy (the sample is imbalanced, seriously affecting its prediction accuracy). Compared with the traditional simple use of transfer learning technology, data augmentation technology, oversampling technology, and data smoothing technology, the present invention can reduce data imbalance and achieve sample expansion of small-sample cities, improve model adaptability and land price prediction accuracy. Adding real samples of similar cities in the present invention can avoid virtual sample expansion by simply using oversampling technology (such as smote), reduce unnecessary noise data, and compared with the disadvantage of data smoothing that important features will be lost in the case of insufficient sample size, the present invention will not destroy the original features of the samples.
[0039] S2. Construct a multi-source index system for land prices, collect multi-source index data of land prices in the research area and cities at the same level as the research area (preferably, the multi-source index data of land prices in the research area and cities at the same level as the research area are data within N3 kilometers around the sample geographical location of the land transaction sample data. In this embodiment, N3 kilometers is exemplified as 5 kilometers) and perform preprocessing. Preferably, the preprocessing of the multi-source index data of land prices includes normalization processing, and the normalization processing expression is as follows: X ij is the data corresponding to index j before normalization processing. i represents the city number of the research area and cities at the same level as the research area. X′ ij is the data corresponding to index j after normalization processing. max(X) is the maximum value of the data before normalization processing. min(X) is the minimum value of the data before normalization processing. The multi-source index data of land prices are derived from point-of-interest data (found by proximity analysis of the geographical locations of the research area and cities at the same level as the research area), economic statistical yearbook data (associated with the regional economic situation), population density data, and remote sensing data. Taking Zhejiang Province as an example, the land transaction sample quantities of each city in Zhejiang Province are as follows in the table:
[0040] Table 1 Table of the Number of Land Transaction Samples in Each City and Region of Zhejiang Province
[0041]
[0042] As can be seen from the above table, Zhoushan is a small-sample city. Based on specific data analysis, although Ningbo has a large number of sample data, there is a problem of data imbalance. In this embodiment, Zhoushan and Ningbo are respectively used as the research areas for implementation examples.
[0043] The multi-source index system of land price includes four category indexes: geographical location, community living facilities, plot attributes, and regional economic conditions. Each category index includes several indexes. In some preferred embodiments, the next level of the geographical location category index includes the shortest distance from the land sample to the hospital, the shortest distance to the hotel, the shortest distance to the factory, the shortest distance to the square, the shortest distance to the villa area, the shortest distance to the commercial service, the shortest distance to the tourist attraction, and / or the shortest distance to the kindergarten. The next level of the community living facilities category index includes the number of commercial services, the number of kindergartens, the number of hotels, the number of primary schools, the number of medical points, the number of bus stops, the number of convenience stores, the number of bars, and / or the number of shopping centers around the land transaction sample. The next level of the plot attributes category index includes population density, normalized difference vegetation index NDVI, minimum floor area ratio, maximum floor area ratio, and / or benchmark land price. The normalized difference vegetation index NDVI is calculated by collecting satellite remote sensing data and extracting the reflectance value NIR of the red light band and the reflectance value Red of the near-infrared band according to the following formula: Regional economic conditions include the per capita living consumption expenditure of residents. Such as Figure 2As shown in the figure, the present invention gives an example of data collection: set the data content items of points of interest, search for points of interest data (including data content items of points of interest such as hospitals, hotels, commercial services, kindergartens, subway stations, bus stops, convenience stores, and / or shopping centers, etc.) within the neighboring range of N3 kilometers around the research area and the geographical locations of cities at the same level as the research area, respectively set the data content items of points of interest for geographical locations and community living facilities and the data of N3 kilometers, calculate the shortest distances from the land samples to hospitals, hotels, factories, squares, villa areas, commercial services, tourist attractions, and / or kindergartens, and construct each index of the geographical location category index, calculate the number of commercial services, kindergartens, hotels, primary schools, medical points, bus stops, convenience stores, bars, and / or shopping centers around the land transaction samples, and construct each index of the community living facilities category index. The acquisition source data of the plot attribute category index includes land transaction cases, remote sensing image data, population density data, and calculate index data such as population density, normalized difference vegetation index NDVI, minimum floor area ratio, maximum floor area ratio, and / or benchmark land price. The acquisition source data of the regional economic status category index includes data such as statistical yearbooks.
[0044] Calculate the average value of all indexes in the multi-source index system of land prices in the research area and cities at the same level as the research area, and screen the top N1 cities at the same level as the research area with the smallest Euclidean distance from the research area as the similar cities of the research area. In some preferred embodiments, the multi-source index data of land prices in the similar cities and the research area are respectively projected to the 2000 National Geodetic Coordinate System by using the Gauss-Kruger projection through ArcMap software and rasterized, and the raster cells are clustered and divided based on the concentration degree of the multi-source index data of land prices, and the position attributes of the raster are correspondingly assigned. The position attribute is the position information of the raster in the raster cell and the similar city or the research area. In the example where Zhoushan and Ningbo are used as the research areas, for the convenience of introducing the technical implementation, the cities at the same level as the research area are only limited to Zhejiang Province in the example implementation of the present invention (in the embodiments of the present invention, it is not necessarily only limited that the research area and the cities at the same level as the research area are in the same province, and this embodiment is only a simple example of the technical implementation). The similar cities of Zhoushan and Ningbo as the research areas are as follows in the table:
[0045] Table 2 Corresponding Similar Cities in Each City of Zhejiang Province
[0046]
[0047] In some embodiments, the calculation expression of the Euclidean distance between the research area and the cities at the same level as the research area is as follows:
[0048] where d is the Euclidean distance, Y 1 、Y2 They are the land transaction prices of the study area and cities at the same level as the study area, X 1n and X 2n They are the average values of the corresponding index n in the study area and cities at the same level as the study area respectively.
[0049] In some embodiments, the land transaction sample data in method S1 includes land price multi-source index data corresponding to the land price multi-source index system. The collection of land price multi-source index data in method S2 includes associative collection from the land transaction sample data and collection of the remaining data that the land transaction sample data does not include land price multi-source index data (of course, it is also possible to collect all multi-source index data including land price multi-source index data, and then perform identification and deduplication processing on the two with the same multi-source index data. When performing deduplication processing, two data of the same multi-source index are verified and verified). The preprocessing of the land price multi-source index data includes data verification processing (constructing a verification relational database, using logically associated data for data verification, removing incorrect data or alarm output of suspicious data); the land transaction price of the land transaction sample data is the land transaction unit price.
[0050] S3. Using the land transaction price of the land transaction sample data as the dependent variable and all the indexes in the land price multi-source index system as the independent variables, construct a land price regression prediction model. Input the similar cities and the study area as two independent databases (if there are multiple similar cities, sub-databases can be established separately in the independent database of similar cities, or merged into a total database) into the land price regression prediction model respectively for separate model training of land transaction price prediction (model training is the regression relationship training between all independent variables and the dependent variable). Obtain the feature SHAP values of the two independent database samples in the land price regression prediction model through the random forest algorithm. Screen the first N2 identical and / or similar ones in the feature SHAP value ranking from the independent database of similar cities to the independent database of the study area as associated similar samples (preferably, the following constraint conditions are also set: the fluctuation between the land transaction unit prices of the cities at the same level as the study area and the associated similar samples of the study area is constrained within M%, and M% is exemplified as 5%). Based on the associated similar samples, use the independent database of similar cities to expand the corresponding data of the independent database of the study area and re-train the model after expansion.
[0051] In the example where the small-sample city Zhoushan is used as the study area, samples similar to the land price of Zhoushan are identified, and samples with a 5% fluctuation between their prices are selected as samples similar to the land price of Zhoushan and used as candidate similar samples. The feature SHAP values of the training samples participated by Zhoushan, Huzhou, Jiaxing, and Lishui can be obtained. The calculation formula of the SHAP value is as follows:
[0052]
[0053] where φ j (x) represents the j-th feature of sample x, F is the set of all features, S is a subset of features, and f s (x s ) represents the predicted value of the model when only the features in subset s are involved, and f s∪{j} represents the predicted value of the model when using the feature subset S∪{j} for prediction, where x s represents the input value of the feature subset S.
[0054] The order of SHAP values of all sample features in a single city (taking Hangzhou as an example) is as Figure 4 shown, and the SHAP value distribution of a single sample (taking Hangzhou as an example) is as Figure 5 shown. Secondly, sort the SHAP values of the features of each sample in Zhoushan and its similar cities by size, and select the top ten features. Then, screen out the samples in which the top three features among the top ten features of the samples in Zhoushan and its similar cities are exactly the same, and among the remaining seven features, at least three features are the common features of the small sample cities to be predicted and the samples of similar cities. The samples that meet the above conditions are identified as similar samples of Zhoushan; the specific expression for determining similar samples is as follows:
[0055]
[0056] where Similarity(N i , N j ) is a group of similar samples, N i is the sample of Zhoushan City, N j is the sample of three similar cities, Huzhou, Jiaxing, and Lishui, and k represents the feature index used to traverse the feature ranking list.
[0057] Using the above method to perform overall expansion within the value range of the samples in Zhoushan City, the number of added samples in Zhoushan City is 759 cases.
[0058] In some embodiments, the SHAP value of the feature of the present invention is obtained through a random forest algorithm model of the SHAP value of an interpretable model. The random forest model is composed of no less than 500 decision trees to form a random forest, and the maximum depth of each decision tree is 15.
[0059] In some embodiments, the screening method for associating similar samples in method S3 is as follows:
[0060] Sort the feature SHAP values of the independent databases in the study area and the independent databases of similar cities respectively, compare the top P1 feature SHAP values in the sorting, and select the first P2 of the independent databases in the study area and the independent databases of similar cities that are exactly the same, and the remaining P1 - P2 are shared by the study area and similar cities as associated similar samples. The example is as follows: P1 is 10, P2 is 3, P1 - P2 is 7. The associated similar samples can be set such that the first 3 feature SHAP values are exactly the same, and the remaining 7 feature SHAP values are shared by the samples of the similar cities and the study area. The samples that meet the condition that the first 3 feature SHAP values are exactly the same and the remaining 7 feature SHAP values are shared by both are used as associated similar samples.
[0061] S4. Input the multi-source index data of land prices in the study area into the land price regression prediction model after model training with capacity expansion and output the predicted land transaction prices. In the example of the small-sample city of Zhoushan as the study area, if only the original data of Zhoushan City is input into the land price regression prediction model without capacity expansion training to obtain the prediction result, and the prediction result is compared with the real result, the comparison chart is as Figure 6 shown, and it can be seen that its accuracy is relatively low. If the method of the present invention expands the sample data of Zhoushan City and then inputs it into the land price regression prediction model after capacity expansion training to obtain the prediction result, and the prediction result is compared with the real result, the comparison chart is as Figure 7 shown, and it can be seen that its accuracy is very high. Compared with Figure 6 the comparison result, the accuracy is greatly improved. Similarly, in the example of the small-sample city of Ningbo as the study area in the present invention, when the prediction result is compared with the real result, the comparison chart is as Figure 8 shown, and it can be seen that its accuracy is relatively low; if the method of the present invention expands the sample data of Ningbo City and then inputs it into the land price regression prediction model after capacity expansion training to obtain the prediction result, and the prediction result is compared with the real result, the comparison chart is as Figure 9 shown, and it can be seen that the data imbalance is reduced and the prediction accuracy is improved.
[0062] If the present invention is extended to some cities in Zhejiang Province and a comparison is made before and after sample expansion according to the land price regression prediction model of the present invention, the precision comparison before and after sample expansion is as follows in the table:
[0063] Table 3 Precision Comparison of Each City in Zhejiang Province Before and After Sample Expansion
[0064]
[0065] As can be seen from the above table, the present invention improves the accuracy for other cities. In particular, it has strong adaptability to small-sample cities (Zhoushan) and sample-imbalanced cities (Ningbo), greatly improving the model prediction accuracy. After the model training of the present invention, the MAPE of the training set increased from the original 0.1204 to 0.1115, and R 2 changed from the original 0.9346 to 0.8914. The MAPE of the test set increased from the original 0.3713 to 0.2881, and R 2 changed from the original 0.1666 to 0.5278. After expanding the land transaction samples in Zhoushan City by the method of the present invention, the R 2 value was significantly improved and greater than 0.5. The meaningless land prediction model obtained from the original data training became somewhat meaningful. For Ningbo City with sufficient sample data but sample imbalance, by the method of the present invention, the imbalance of the data was reduced. After model training, the MAPE of its training set increased from the original 0.0774 to 0.0772, and R 2 changed from the original 0.9749 to 0.9746. The MAPE of the test set increased from the original 0.1921 to 0.792, and R 2 changed from the original 0.8368 to 0.8684. The R 2 of the test sets of other cities (except Jinhua, Lishui, and Shaoxing) all increased to a certain extent. The MAPE of all cities increased, and the overall average of the model prediction accuracy before and after sample expansion increased by 4.30%, and R 2 increased by 8.97% on average overall. The practical results show that the sample expansion method for predicting land prices in small-sample cities proposed by the present invention is applicable to most cities in Zhejiang Province, solves the problem of insufficient sample data volume, optimizes the problem of data imbalance, improves the generalization ability and practical application effect of the model, enhances the data utilization efficiency and the stability of the model, and provides a replicable solution, improving the practical application value of land price prediction.
[0066] A land price prediction system for small-sample urban data expansion, including a data acquisition and processing module, a similar city database, a research area database, a multi-source index system for land prices, a similar city screening module, a land price regression prediction model, and an output module. The multi-source index system for land prices includes four category indexes: geographical location, community living facilities, plot attributes, and regional economic conditions. Each category index includes several indexes. The data acquisition and processing module obtains the land transaction sample data and multi-source index data of land prices in the research area, preprocesses them, and stores them in the research area database. The data acquisition and processing module obtains the land transaction sample data and multi-source index data of land prices in the same-level cities of the research area, preprocesses them, and stores them in the database of the same-level cities of the research area. The land transaction sample data includes land transaction prices, and the research area is a small-sample city. The similar city screening module calculates the average value of all indexes in the multi-source index system of land prices in the research area and the same-level cities of the research area, and screens the top N1 same-level cities of the research area with the smallest Euclidean distance from the research area as the similar cities of the research area and stores them in the similar city database. The similar city screening module is constructed with the land transaction price of the land transaction sample data as the dependent variable and all indexes in the multi-source index system of land prices as the independent variables. The similar cities and the research area are used as two independent databases and are respectively input into the land price regression prediction model for separate model training of land transaction price prediction. The feature SHAP values of the two independent database samples in the land price regression prediction model are obtained through the random forest algorithm model. The top N2 same and / or similar ones in the feature SHAP value ranking of the independent database of the similar cities are screened from the independent database of the similar cities as the associated similar samples. Based on the associated similar samples, the independent database of the similar cities is used to perform corresponding data expansion on the independent database of the research area and re-model training after the expansion. The multi-source index data of land prices in the research area is collected and input into the land price regression prediction model after model training after the expansion, and the predicted land transaction price is output. The output module is used to output the predicted land transaction price of the research area.
[0067] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A land price prediction method based on small sample city data expansion, characterized by: The methods include: S1. Take a small sample city as the study area, obtain the land transaction sample data of the study area and cities of the same level as the study area, and preprocess the data. The land transaction sample data includes land transaction prices; S2. Construct a multi-source indicator system for land prices. Collect and pre-process the multi-source indicator data for land prices in the study area and cities of the same level in the study area. The multi-source indicator system for land prices includes four categories of indicators: geographical location, community living facilities, land plot attributes, and regional economic conditions. Each category of indicators includes several indicators. Calculate the average values of all indicators in the multi-source indicator system for land prices in the study area and cities of the same level in the study area, and select the first N1 cities of the same level in the study area with the smallest Euclidean distance to the study area as similar cities in the study area. S3. A land price regression prediction model is constructed with the land transaction price of the land transaction sample data as the dependent variable and all the indicators in the land price multi-source indicator system as the independent variables. Similar cities and the study area are respectively input into the land price regression prediction model as two independent databases to perform separate model training for land transaction price prediction. The characteristic SHAP values of the two independent database samples in the land price regression prediction model are obtained by the random forest algorithm. The first N2 identical and / or similar samples in the characteristic SHAP value ranking are selected from the independent database of similar cities and the independent database of the study area as the associated similar samples. Based on the associated similar samples, the independent database of similar cities is used to expand the corresponding data of the independent database of the study area and retrain the model after the expansion. S4. Collect multi-source indicator data of land prices in the study area and input them into the land price regression prediction model after the expanded model training and output the predicted land transaction price.
2. The land price prediction method based on small sample city data expansion according to claim 1 is characterized by: In method S1, the land transaction sample data includes land price multi-source indicator data corresponding to the land price multi-source indicator system, and in method S2, collecting the land price multi-source indicator data includes performing associated collection from the land transaction sample data and collecting the remaining data that does not include the land price multi-source indicator data in the land transaction sample data, and the preprocessing of the land price multi-source indicator data includes data verification processing; The land transaction price in the land transaction sample data is the land transaction unit price.
3. The land price prediction method based on small sample city data expansion according to claim 1 is characterized by: The next level of the geographical location category index includes the shortest distance from the land sample to the hospital, the shortest distance from the hotel, the shortest distance from the factory, the shortest distance from the square, the shortest distance from the villa area, the shortest distance from the commercial service, the shortest distance from the tourist attraction or / and the shortest distance from the kindergarten; the next level of the community life facility category index includes the number of commercial services, kindergartens, hotels, primary schools, medical points, bus stations, convenience stores, bars or / and shopping centers around the land transaction sample; the next level of the land plot attribute category index includes population density, normalized vegetation index NDVI, minimum plot ratio, maximum plot ratio or / and benchmark land price. The normalized vegetation index NDVI is calculated by collecting satellite remote sensing data and extracting the red light band reflectance value NIR and the near infrared band reflectance value Red according to the following formula: The economic conditions of the region include per capita living consumption expenditure of residents.
4. The land price prediction method based on small sample city data expansion according to claim 1 or 3, characterized in that: The multi-source indicator data of land prices in similar cities and study areas were projected to the 2000 National Geodetic Coordinate System using the Gauss Kruger projection through ArcMap software and rasterized. The grid units were clustered based on the concentration of the multi-source indicator data of land prices and the location attributes of the grids were assigned accordingly. The location attributes are the location information of the grid in the grid unit and similar cities or study areas.
5. The land price prediction method based on small sample city data expansion according to claim 1 is characterized by: In method S2, the multi-source indicator data of land prices in the study area and cities of the same level in the study area are data within N3 kilometers around the sample geographical location of the land transaction sample data.
6. The land price prediction method based on small sample city data expansion according to claim 1 is characterized by: The calculation expression of the Euclidean distance between the study area and the cities of the same level in the study area is as follows: Where d is the Euclidean distance, Y1 and Y2 are the land transaction prices of the study area and the cities of the same level in the study area, respectively. 1n , X 2n They are the average values of the corresponding index n in the study area and cities of the same level in the study area, respectively.
7. The land price prediction method based on small sample city data expansion according to claim 1 is characterized by: The preprocessing of multi-source land price index data in method S2 includes normalization processing, and the normalization processing expression is as follows: X ij is the corresponding data of index j before normalization; i represents the city number of the study area and the cities of the same level in the study area; X′ ij is the data corresponding to index j after normalization; max(X) is the maximum value of the data before normalization; min(X) is the minimum value of the data before normalization.
8. The land price prediction method based on small sample city data expansion according to claim 1 is characterized by: The feature SHAP value is obtained through the random forest algorithm model that can interpret the model SHAP value. The random forest model uses no less than 500 decision trees to form a random forest, and the maximum depth of each decision tree is 15.
9. The land price prediction method based on small sample city data expansion according to claim 1 is characterized by: The screening method for related similar samples in method S3 is as follows: The feature SHAP values of the independent database of the study area and the independent database of similar cities are sorted respectively, and the top P1 feature SHAP values are compared. The independent database of the study area and the independent database of similar cities whose first P2 are exactly the same and the remaining P1-P2 are shared by the study area and similar cities are selected as associated similar samples.
10. A land price prediction system with small sample city data expansion, characterized by: It includes a data acquisition and processing module, a similar city database, a study area database, a land price multi-source indicator system, a similar city screening module, a land price regression prediction model and an output module. The land price multi-source indicator system includes four category indicators: geographical location, community living facilities, plot attributes and regional economic conditions. Each category indicator includes several indicators. The data acquisition and processing module obtains land transaction sample data and land price multi-source indicator data of the study area and stores them in the study area database after pre-processing. The data acquisition and processing module obtains land transaction sample data and land price multi-source indicator data of cities of the same level in the study area and stores them in the database of cities of the same level in the study area after pre-processing. The land transaction sample data includes land transaction prices, and the study area is a small sample city. The similar city screening module calculates the average values of all indicators in the land price multi-source indicator system of the study area and cities of the same level in the study area and screens the top N1 cities of the same level in the study area with the smallest Euclidean distance to the study area as similar cities of the study area. The cities are stored in the similar city database; the similar city screening module is constructed with the land transaction price of the land transaction sample data as the dependent variable and all the indicators in the land price multi-source indicator system as the independent variables, and the similar cities and the study area are respectively input into the land price regression prediction model as two independent databases to perform separate model training for land transaction price prediction; the feature 5HAP values of the two independent database samples in the land price regression prediction model are obtained through the random forest algorithm model, and the independent databases of the similar cities are screened as the first N2 identical and / or similar to the independent database of the study area in the feature SHAP value sorting as the associated similar samples, and the independent database of the similar cities is used to expand the corresponding data of the independent database of the study area based on the associated similar samples, and the model is retrained after the expansion; the multi-source indicator data of the land price of the study area is collected and input into the land price regression prediction model after the expansion model training, and the predicted land transaction price is output, and the output module is used to output the predicted land transaction price of the study area.