Multi-agent network big data city house price analysis and calculation method
Through multi-intelligent network big data analysis and combining multiple data sources to build a city's second-hand housing price impact factor model, the problem of thin housing price calculation model in the existing technology is solved, and high-precision housing price distribution law analysis and prediction are achieved.
Patent Information
- Application Number
- CN202510321916.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology fails to effectively consider multi-factor spatial data in urban second-hand housing prices analysis, lacks spatial heterogeneity analysis and autocorrelation analysis, and cannot accurately model location and environmental factors, resulting in the housing price calculation model being too thin, with large errors and low degree of refinement.
Through multi-agent network big data analysis, combined with POIs data, road network data and Landsat data, a city second-hand housing price impact factor model is constructed, a spatial heterogeneity and autocorrelation analysis of housing price space is carried out, and a zoning nonlinear characteristic price model is established to quantify the impact of commercial development, transportation, infrastructure, location, education, environment and residents' consumption level.
A high-precision urban second-hand housing price distribution pattern model is realized, housing prices are calculated in a refined manner, and site selection suggestions for high-end communities are provided, which improves the accuracy and accuracy of housing price prediction.
Smart Images

Figure CN120471636A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a method for calculating city housing prices based on network big data, and in particular to a method for analyzing and calculating city housing prices based on multi-agent network big data, belonging to the technical field of big data housing price calculation. Background Art
[0002] Real estate is a high value-added industry in the national economy. The relationship between the specific layout of housing prices in cities and the surrounding locational environmental factors, and how to reasonably regulate the housing market through macro-control methods have become urgent issues that need to be addressed and are also issues of general concern. Revealing the laws of housing prices has positive significance for urban planning and real estate pricing.
[0003] Existing technologies have studied urban housing prices from various perspectives. Urban housing prices are closely related to their spatial location and surrounding infrastructure, resulting in different housing prices depending on the location. The competitive rent function, based on the single-center assumption, posits that urban housing rents decrease with increasing distance from the city center. However, due to the diversity and complexity of urban systems, the emergence of secondary urban centers, and improved transportation, simply using distance to the city center as a parameter for evaluating housing price distribution is insufficient. A more accurate housing price calculation model is urgently needed by integrating multiple social, economic, locational, and environmental factors. This is crucial for studying the distribution patterns of urban housing prices.
[0004] The problems that need to be solved by the existing urban second-hand housing price analysis and calculation methods and the key technical difficulties of this application include:
[0005] (1) The bidding function of the existing technology is based on the single-center assumption and concludes that the urban housing rent is a negative correlation function that decreases with distance from the city center to the periphery. However, due to the diversity and complexity of the urban system, the emergence of urban secondary centers, and the improvement of traffic conditions, the model that simply uses the distance to the city center as a parameter to evaluate the distribution of housing prices is too thin. The existing technology has not analyzed the multi-factor spatial data of urban second-hand housing prices, lacks the analysis of spatial heterogeneity of housing prices and the analysis of spatial autocorrelation of housing prices, lacks the modeling of factors affecting urban second-hand housing prices, does not construct the factors affecting urban second-hand housing prices and the quantification of the importance of factors affecting factors based on multi-agents, and lacks multiple aspects to measure the location environment factors that affect housing prices; does not establish a model for the spatial distribution law of urban second-hand housing prices, lacks the urban second-hand housing characteristic price regression model and the second-hand housing price geographical weighted regression model, and does not jointly calculate and analyze the spatial distribution law of housing prices. At present, it is urgent to construct characteristics from multiple factors such as social, economic, location, and environmental factors to make the housing price calculation model more accurate, which plays an important role in calculating the distribution law of urban housing prices.
[0006] (2) The existing technology lacks a reliable method for analyzing the multi-factor spatial data of urban second-hand housing prices, lacks the analysis of original housing price data based on partition calculation and histogram calculation, and cannot determine the basic spatial layout and calculation characteristics of urban second-hand housing prices; lacks the spatial heterogeneity analysis of second-hand housing prices in various administrative districts of the city based on geographic detectors, and fails to obtain the distribution differences of housing prices between different administrative districts; lacks the global spatial autocorrelation analysis of urban second-hand housing prices based on Moran's I, and cannot obtain the spatial autocorrelation and spatial aggregation characteristics of urban second-hand housing prices in space. There are few factors considered in second-hand housing prices, and the spatial data analysis is not reliable enough. The prediction of second-hand housing prices has a large error compared with the actual situation.
[0007] (3) The existing technology lacks an accurate modeling method for factors affecting urban second-hand housing prices. It has not established a model for the impact of locational environmental factors on housing prices. It has not used POIs data, review POIs data, road network data, and Landsat data to quantify the distribution of various social, economic, and locational factors such as commercial development, transportation, infrastructure, location, education, environment, and residents' consumption levels near the community. It is impossible to obtain the contribution of various characteristics to the spatial variation of housing prices in nonlinear models.
[0008] (4) The existing technology lacks a model for the spatial distribution of urban second-hand housing prices, and lacks a method to establish different characteristic price models based on the results of multi-factor spatial analysis and influencing factor analysis to obtain the results of urban housing price distribution. It is unable to further analyze the spatial layout of the urban housing price market and distinguish between high-value areas, transition areas and low-value areas. The existing technology lacks a high-precision spatial distribution model for urban second-hand housing prices and a model for the analysis of factors affecting housing prices. It does not realize the use of plots divided by road networks as the smallest unit to represent the urban second-hand housing price object. Since the housing itself or other abnormal factors have a greater impact on housing prices, the urban housing price analysis and calculation results obtained do not fully combine various factors, have large errors, and are of low refinement. Summary of the Invention
[0009] This application analyzes the multi-factor spatial data of urban second-hand housing prices, and adopts housing price spatial heterogeneity analysis and housing price spatial autocorrelation analysis in sequence, and creatively proposes modeling of factors affecting urban second-hand housing prices, including constructing factors affecting urban second-hand housing prices and quantifying the importance of factors based on multi-agents, and constructing factors affecting urban second-hand housing prices to measure the locational environmental factors that affect housing prices from seven aspects: business development, infrastructure, transportation, location, education, consumption level, and environment; creatively establishes a model for the spatial distribution law of urban second-hand housing prices, including an urban second-hand housing characteristic price regression model and a second-hand housing price geographically weighted regression model, and jointly calculates and analyzes the spatial distribution law of housing prices, fully combining the weights of different locational environmental factors to calculate the characteristics of urban housing prices in different detailed locations, and fully combining the spatiotemporal characteristics to efficiently and accurately calculate second-hand housing prices.
[0010] To achieve the above technical effects, the technical solutions adopted in this application are as follows:
[0011] A multi-agent network big data urban housing price analysis and calculation method: First, the multi-factor spatial data of urban second-hand housing prices is analyzed, and housing price spatial heterogeneity analysis and housing price spatial autocorrelation analysis are used in sequence. Then, the influencing factors of urban second-hand housing prices are modeled, including the construction of urban second-hand housing price influencing factors and the quantification of the importance of influencing factors based on multi-agents. The construction of urban second-hand housing price influencing factors measures the locational environmental factors that affect housing prices from seven aspects: commercial development, infrastructure, transportation, location, education, consumption level, and environment. Finally, a model for the spatial distribution of urban second-hand housing prices is established, including an urban second-hand housing characteristic price regression model and a second-hand housing price geographically weighted regression model, and the spatial distribution of housing prices is calculated and analyzed in combination.
[0012] 1) Analyzing multi-factor spatial data on urban second-hand housing prices: Using zoning and histogram calculations, we analyzed the original housing price data to determine the basic spatial distribution and calculation characteristics of urban second-hand housing prices. We also used geographic detectors to analyze the spatial heterogeneity of second-hand housing prices in various administrative districts of the city, determining the distribution differences between housing prices in different districts. We also used Moran's I tool to conduct a global spatial autocorrelation analysis of urban second-hand housing prices, determining the spatial autocorrelation and spatial clustering characteristics of urban second-hand housing prices.
[0013] 2) Modeling factors influencing urban second-hand housing prices: Build a model for the impact of locational environmental factors on housing prices. Utilize POI data, review POI data, road network data, and Landsat data to quantify the distribution of various social, economic, and locational factors near residential areas, including commercial development, transportation, infrastructure, location, education, environment, and resident consumption levels. These characteristics are then integrated into a multi-intelligence assessment model to determine the contribution of various characteristics to the spatial variation of housing prices in a nonlinear model.
[0014] 3) Spatial distribution model of urban second-hand housing prices: Based on the results of multi-factor spatial analysis and influencing factor analysis, different characteristic price models are established to obtain high-precision urban housing price distribution results. Among them, the partitioned nonlinear characteristic price model has higher accuracy than the linear model and the global modeling model. The local Moran's I is further used to analyze the spatial layout of the urban housing price market and distinguish high-value areas, transition areas, and low-value areas.
[0015] Preferably, the spatial heterogeneity analysis of housing prices is performed by calculating the difference in housing prices in different administrative areas based on geographic detectors. Under the condition of discrete independent variables, the dependent variable is divided into blocks. The difference between different blocks is the spatial heterogeneity. The q value of the differentiation detector detects the spatial heterogeneity of the dependent variable. The expression is:
[0016]
[0017] Where L is the stratification of factor X, N h and N are the number of samples in each layer and the whole area, and σ 2 is the variance of the dependent variable of each district and the whole district, and the value range of q is [0, 1]. The larger the value, the greater the difference of the dependent variable among the districts.
[0018] Analysis of spatial autocorrelation of housing prices: Based on the spatial clustering of housing price datasets, the spatial autocorrelation is calculated, which is expressed as follows:
[0019]
[0020] where z i is the difference between the attribute value of sample i and the sample mean, w i,j is the spatial weight between sample i and sample j, n is the number of samples, and when the value of I is 0, the housing price has nothing to do with its spatial position, and the housing price is randomly distributed in space; when the value is closer to 1, it indicates that there are high values and high value clusters of housing prices in space. On the contrary, the closer it is to -1, it indicates that there is a step-like fault in the distribution of housing prices, and there are abnormal land price areas around high-priced areas.
[0021] Preferably, infrastructure factor: medical POIs are used to describe the distribution of medical resources, including: first-level hospitals, second-level hospitals, third-level hospitals, and private clinics;
[0022] The distribution of nearby infrastructure is measured by calculating the number of life service POIs within 1500m of the cell, i.e., the point density.
[0023] Construct infrastructure factor evaluation indicators. The evaluation formula is:
[0024] F2(x, y)=I6Dis6(x, y)+I7Den7(x, y) Formula 4
[0025] Among them, F2(x, y) is the infrastructure factor evaluation index, Dis6 and Den7 are the distance to medical POIs and the density of life service POIs calculated using POIs data, respectively. i is the coefficient factor of each feature.
[0026] Optimally, transportation factors: POIs are captured using transportation facilities as keywords, with subway stations and bus stations as intra-city transportation stations, and airports, railway stations, port terminals, and long-distance bus stations as inter-city transportation stations to analyze the impact of different transportation facilities on housing prices;
[0027] The road network structure and road network-level data reflect the city's transportation network from a physical perspective and are closely related to the distribution of housing prices within the city. The road network data in this application includes national highways, provincial highways, expressways, urban main roads, urban secondary roads, urban branch roads, miscellaneous roads, rural roads, and county roads. By calculating the distance to each level of road, the road network structure near the community is also measured. By integrating these multiple factors, a traffic factor evaluation index is constructed. The evaluation formula is as follows:
[0028]
[0029] Among them, F3(x, y) is the traffic factor evaluation index, Dis8 is the distance to each transportation facility calculated using POIs data, and Dis i , i = 9, ..., 17 is the distance to national roads, provincial roads, expressways, urban main roads, urban secondary roads, urban branch roads, miscellaneous roads, rural roads, and county roads. i is the coefficient factor of each feature.
[0030] Preferably, the location factor is a price-influencing factor of the community's location in the city. By capturing POIs of central government agencies, public security, procuratorial and judicial agencies, and government agencies at all levels, and calculating their point density as one of the location conditions for measuring the community, a location factor evaluation index is constructed. The evaluation formula is as follows:
[0031] F4(x, y)=1 18 Dis 18 (x, y) + I 19 Den 19 (x, y) Formula 6
[0032] Among them, F2(x, y) is the location factor evaluation index, Dis 18 、Den 19 They are the distance to the CBD of the convention and exhibition center and the density of government POIs, I i is the coefficient factor of each feature.
[0033] Preferably, the education factor is: by collecting the public high school enrollment rates of each school and the schools published on the education website as a quantitative indicator of the school's teaching staff. In terms of school district housing, by collecting the admission rules of each middle school provided by the government online, excluding schools that serve the entire region, the enrollment districts divided by each school are recorded, and the location query service provided by the map is used to obtain the corresponding school district location information to construct an education factor evaluation index. The evaluation formula is as follows:
[0034] F5(x, y)=1 20 P 20 (x, y) Formula 7
[0035] Among them, F5(x, y) is the evaluation index of education factor, P 20 (x, y) is the enrollment rate of junior high schools in the school district where the community is located, I i is the coefficient factor of each feature.
[0036] Preferably, the consumption level factor is based on the Dianping POIs data, and the average per capita consumption of restaurants near the community is obtained by performing the distance inverse weighted interpolation method on the food consumption point data as a measure of the consumption level of the community population, and a consumption level factor evaluation index is constructed. The evaluation formula is as follows:
[0037] F6(x, y)=1 21 P 21 (x, y) Formula 8
[0038] Among them, F6(x, y) is the consumption level factor evaluation index, P 21 (x, y) is the average per capita consumption of restaurants near the community, I i is the coefficient factor of each feature.
[0039] (7) Environmental factors
[0040] Based on Landsat8 data, the Normalized Difference Vegetation Index (NDVI) and the Modified Normalized Difference Water Index (MNDWI) were calculated. The average NDVI and MNDWI of each plot were calculated to obtain the vegetation and water cover of the community. The calculation formulas for NDVI and MNDWI are as follows:
[0041]
[0042] Among them, MIR is the mid-infrared band, NIR is the near-infrared band, R is the red band, and G is the green band;
[0043] Construct environmental factor evaluation indicators, and the evaluation formula is as follows:
[0044] F7(x, y)=1 22 NDVI+I 23 MNDWI formula 10
[0045] Among them, F7(x, y) is the environmental factor evaluation index, NDVI and MNDWI are the normalized vegetation index and improved normalized water index, I i is the coefficient factor of each feature.
[0046] Preferably, the importance of influencing factors is quantified based on multi-agents: there is a nonlinear functional relationship between housing prices and location and environmental factors. The importance of each factor in the nonlinear model of housing prices is analyzed by using the feature importance evaluation in the multi-agent. The multi-agent model is an integrated model, in which the agent is a tree-like prediction model. The training sample data set is input into the agent, and the sample is segmented by selecting some features. In each segmentation, a feature is selected for critical value division. Finally, the samples of different leaf nodes have different category attributes, and the agent structure is adjusted by pruning.
[0047] In the feature selection of generating agents, input variables are randomly selected, and part of the features of the input sample set are used to construct the agent. The selection process is also sampling with replacement, and a tree-structured data partition is constructed. By constructing agents multiple times, a multi-agent model is obtained. Through multiple random processes, the multi-agent model can better analyze the distribution of the overall data than a single agent. A multi-agent is a system composed of a series of agents h k (x)(k=1,…,N), the marginal function is defined as:
[0048] mg(X, Y)=av k (I(h k (X)=Y))-max j≠Y (av k (I(h k (X)=j))) Formula 11
[0049] Where I(·) is the indicator function, Y is the correct classification vector, j is the incorrect classification vector, and av k (·) means taking the average, and the marginal function indicates the degree to which the number of votes for the correct classification exceeds the maximum number of votes for the incorrect classification. The larger the value, the higher the confidence in the classifier;
[0050] The generalization error is used to show how well the model fits the training data, to prevent overfitting caused by the inability of the training data to estimate the distribution of the entire data and the model being overtrained to fit the training data. It is defined as:
[0051] PE=P X,Y (mg(X, Y)<0) Formula 12
[0052] Among them, P X,Y (·) indicates that the generalization error is obtained under the distribution of random variables X and Y;
[0053] The generalization error PE will tend to an upper bound, and the multi-agent algorithm will not suffer from serious overfitting problems as the number of agents increases. One of the known upper bounds of the generalization error is:
[0054]
[0055] Where avp is the average correlation between agents, s is the average strength of agents. To achieve good generalization performance of multi-agents, the correlation between agents should be reduced and the average strength of agents should be increased.
[0056] The grid search method is used for search. After obtaining the optimal parameters, the feature importance of the feature set is calculated to evaluate the impact of all features on the price of second-hand houses in the city. The feature importance is calculated multiple times and the ones with out-of-bag R-square less than 0 are removed. The 10 models are weighted according to their out-of-bag R-square to obtain the average feature importance.
[0057] Optimally, the urban second-hand housing hedonic price regression model: uses hedonic price regression to explore the relationship between the built environment and second-hand housing prices. After model training, it can use POI density feature information to calculate housing price information; and fits the nonlinear function relationship by minimizing empirical risk.
[0058] There is a training sample set (X, Y) = {x 1i ,…,x si ,y i}, i = 1, ..., N, where N is the number of samples, X = {x 1i ,…,x si} is the sample feature set, s is the number of features, y i For the corresponding output value, the relationship is:
[0059]
[0060] in is a kernel function that maps the X vector to a high-dimensional space and makes it have a linear relationship with the output value Y: the ω and b parameters are solved using the following formula:
[0061]
[0062] where ξ i and is a relaxation factor that reduces the noise in the data to avoid overfitting; ∈ is the bound of the loss function, which allows a reasonable error between the regression result and the true value; C is a parameter that balances the regression result and the relaxation variable;
[0063] Solve the constrained quadratic programming problem and find ω and b as follows:
[0064]
[0065] where α i , is the minimization solution of J(ω, e), and finally the model expression is obtained:
[0066]
[0067] in is the kernel function.
[0068] Preferably, the geographically weighted regression model for second-hand housing prices is used: by inputting the coordinates of the variables into the model as parameters, and using non-parametric estimation to obtain the corresponding local function estimate at each location, the distribution of the dependent variable is analyzed in detail. Based on the first law of geography, different regions are modeled separately to reflect the spatial heterogeneity of the distribution of housing prices and influencing factors, as well as the spatial heterogeneity of the relationship between housing prices and influencing factors. The model is expressed as:
[0069]
[0070] Where, (u i , v i ) is the geographical coordinate of the center point of the i-th road network segmentation object, x ik is the characteristic variable of the i-th object, β0(u i , v i ) is the sum of the constants of other influencing factors of the characteristic variable, β k (u i , v i ) is the feature x ik The regression coefficient, ε i is a random error term; a linear function is used to characterize the characteristic regression coefficient of housing prices, and the expression is:
[0071] β t =λ t0 +λ t1 u i +λ t2 v i Formula 21
[0072] For any point in the region (u i , v i ) is estimated by least squares:
[0073]
[0074] Where X is the impact factor matrix, W(u i , v i ) is a weight matrix, which is composed of the monotonically decreasing function value of the spatial position distance between the regression object and other surrounding objects. The kernel function expression is as follows:
[0075]
[0076] Where d is the Euclidean distance between objects and h is the optimal bandwidth, which is obtained by the cross determination method:
[0077]
[0078] in is the value of the house obtained after fitting the i-th object, y i is the true value of the house price at object i.
[0079] Compared with the existing technology, the innovation and advantages of this application are:
[0080] (1) This application uses a crawler to capture urban second-hand housing data and Baidu point of interest data, and obtains the computational characteristics of urban housing prices by performing computational analysis on the original housing price data, as well as spatial autocorrelation analysis and spatial heterogeneity analysis. It further combines the micro-level housing price influencing factors, including commercial development, transportation, infrastructure, location, education, environment and residents' consumption level, and uses a crawler to capture Baidu POIs data, Dianping POIs data, Landsat8 data, and education network data to quantify the influencing factors. It combines the captured housing price data to perform feature importance analysis, uses this feature set to perform housing price regression analysis, uses the computational characteristics of urban housing prices and each feature to construct a partitioned nonlinear characteristic price model, and obtains precise and accurate urban second-hand housing price results. It also uses the local Moran index to analyze the spatial structure of the urban housing price market, specifically for the precise calculation of housing prices for specific real estate projects, which has great application value.
[0081] (2) This application obtains network big data by crawling, and constructs features based on this to quantify the factors affecting housing prices, such as commercial development, transportation, infrastructure, location, education, environment and residents' consumption level, and uses a multi-agent nonlinear feature importance evaluation model to quantify the degree of influence of housing price image factors; constructs a partitioned nonlinear feature price model based on spatial heterogeneity, and proposes that different housing price distribution patterns and calculation methods exist in different parts of the city. The superiority of the model is confirmed by experimental comparison, and a local spatial autocorrelation analysis is performed on the model results to obtain a detailed spatial structure of the urban housing price market. Further analysis of the abnormal results provides suggestions for urban renewal and the site selection of high-end residential communities for real estate developers, and the refined calculation of second-hand housing prices is efficient and accurate.
[0082] (3) This application analyzes the multi-factor spatial data of urban second-hand housing prices, and adopts housing price spatial heterogeneity analysis and housing price spatial autocorrelation analysis in turn, and creatively proposes a modeling of factors affecting urban second-hand housing prices, including the construction of urban second-hand housing price influencing factors and the quantification of the importance of influencing factors based on multi-agents, and the construction of urban second-hand housing price influencing factors to measure the locational environmental factors that affect housing prices from seven aspects: business development, infrastructure, transportation, location, education, consumption level, and environment; creatively establishes a model of the spatial distribution law of urban second-hand housing prices, including an urban second-hand housing characteristic price regression model and a second-hand housing price geographically weighted regression model, and jointly calculates and analyzes the spatial distribution law of housing prices, fully combines the weights of different locational environmental factors to calculate the characteristics of urban housing prices in different detailed locations, and fully combines the spatiotemporal characteristics to efficiently and accurately calculate the second-hand housing prices. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 This is a schematic diagram of the locational environmental factors that affect housing prices.
[0084] Figure 2 This is a chart showing the calculation results of educational factors and educational data of a certain city.
[0085] Figure 3 This is a chart showing the corrected normalized water index calculation results for a certain city.
[0086] Figure 4 It is a schematic diagram of the quantification of the importance of influencing factor features based on multiple agents. DETAILED DESCRIPTION
[0087] Below, in conjunction with the accompanying drawings, the technical solution of the multi-agent network big data urban housing price analysis and calculation method provided by this application is further described so that technical personnel in this field can better understand this application and implement it.
[0088] This application, based on the collection of urban second-hand housing transaction prices and urban location environment information, establishes the spatial distribution pattern of urban second-hand housing prices and an analysis model of housing price influencing factors. It uses plots divided by road networks as the smallest unit to represent the community for urban second-hand housing prices, which can alleviate the impact of the house itself or other abnormal factors on housing prices.
[0089] Multi-factor spatial data analysis of urban second-hand housing prices: Based on zoning calculation and histogram calculation, the original housing price data was analyzed to obtain the basic spatial layout and calculation characteristics of urban second-hand housing prices; based on the geographic detector, the spatial heterogeneity of second-hand housing prices in each administrative district of the city was analyzed to obtain the distribution differences of housing prices between different administrative districts; based on the Moran's I tool, the global spatial autocorrelation analysis of urban second-hand housing prices was conducted to obtain the spatial autocorrelation and spatial aggregation characteristics of urban second-hand housing prices. This conclusion illustrates the feasibility of housing price modeling.
[0090] Modeling factors influencing urban second-hand housing prices: We established a model for the impact of locational environmental factors on housing prices. We used POI data, Dianping POI data, road network data, and Landsat 8 data to quantify the distribution of various social, economic, and locational factors near residential communities, including commercial development, transportation, infrastructure, location, education, environment, and residents' consumption levels. We then integrated these characteristics into a multi-intelligence assessment model to determine the contribution of various factors to the spatial variation of housing prices in a nonlinear model. The most important factors were education, transportation, consumption level, and location, followed by commercial development, infrastructure, and environmental factors.
[0091] Spatial distribution model of urban second-hand housing prices: Based on the results of multi-factor spatial analysis and influencing factor analysis, different characteristic price models are established to obtain high-precision urban housing price distribution results. Among them, the partitioned nonlinear characteristic price model has higher accuracy than the linear model and the global modeling model. The local Moran's I is further used to analyze the spatial layout of the urban housing price market, distinguishing high-value areas, transition areas and low-value areas. The model calculates the urban villages in the central area of the city and the high-end residential communities in the elegant suburban areas.
[0092] 1. Analysis of multi-factor spatial data on urban second-hand housing prices
[0093] (1) Analysis of spatial heterogeneity of housing prices
[0094] Based on the geographic detector, the housing price difference of different administrative areas is calculated. Under the condition of discrete independent variables, the dependent variable is divided into blocks. The difference between different blocks is spatial heterogeneity. The q value of the differentiation detector detects the spatial heterogeneity of the dependent variable. The expression is:
[0095]
[0096] Where L is the stratification of factor X, N h and N are the number of samples in each layer and the whole area, and σ 2 is the variance of the dependent variable of each district and the whole district. The value range of q is [0, 1]. The larger the value, the greater the difference of the dependent variable among the districts.
[0097] (2) Analysis of spatial autocorrelation of housing prices
[0098] Based on the spatial aggregation of the housing price dataset, the spatial autocorrelation is calculated, which is expressed as follows:
[0099]
[0100] where z iis the difference between the attribute value of sample i and the sample mean, w i,j is the spatial weight between sample i and sample j, n is the number of samples, and when the value of I is 0, the housing price has nothing to do with its spatial position, and the housing price is randomly distributed in space; when the value is closer to 1, it indicates that there are high values and high value clusters of housing prices in space. On the contrary, the closer it is to -1, it indicates that there is a step-like fault in the distribution of housing prices, and there are abnormal land price areas around high-priced areas.
[0101] 2. Modeling of factors affecting urban second-hand housing prices
[0102] A housing price characteristic model is constructed based on seven aspects, including commercial development, transportation, infrastructure, location, education, environment and residents' consumption level. These influencing factors are quantified and combined with the characteristic importance model to evaluate the impact of different characteristics on urban housing prices.
[0103] (1) Constructing factors affecting urban second-hand housing prices
[0104] like Figure 1 , measuring the locational environmental factors that affect housing prices from seven aspects: commercial development, infrastructure, transportation, location, education, consumption level, and environment.
[0105] (1) Business development factors
[0106] Obtain spatial location data of points of interest (POIs) with social attributes, including name, category, latitude and longitude, and address elements, to display the distribution of facilities and human activities in the city from a geographical perspective. Set shopping, food, leisure and entertainment, finance, and corporate categories in POIs to reflect the impact of commercial development on housing prices.
[0107] Shopping centers are used as an evaluation factor for the degree of commercial development in a region. The better the commercial development, the higher the land price and the higher the housing prices nearby. Based on shopping POI data, the number of POIs within 1500 meters of each plot is calculated, that is, the shopping POI point density to construct the impact factor;
[0108] Based on the restaurant POI data, the number of POIs within 1500 meters of each plot is calculated, that is, the restaurant POI point density to construct the impact factor;
[0109] Based on the leisure and entertainment POI data, the number of POIs within 1500 meters of each plot is calculated, that is, the leisure and entertainment POI point density to construct the impact factor;
[0110] Based on the financial POI data, the number of POIs within 1500 meters of each plot is calculated, that is, the density of financial POIs to construct the impact factor;
[0111] Using "company" as the keyword, we crawled POIs to identify firms, industrial and commercial areas, real estate development, cultural media, travel agencies, property management, telecommunications companies, and high-tech parks to reflect the distribution of companies. Based on the company-related POI data, we calculated the number of POIs within 1,500 meters of each plot, i.e., the company-related POI point density, to construct the impact factor.
[0112] Through multi-agents, we construct an evaluation index for business development factors. The evaluation formula is:
[0113] F1(x,y)=I1Den1(x,y)+I2Den2(x,y)+I3Den3(x,y)
[0114] +I4Den4(x, y)+I5Den5(x, y) Equation 3
[0115] Where F1(x, y) is the comprehensive commercial development factor of a certain place in the city, Den1, Den2, Den3, Den4, and Den5 are the point densities of shopping, food, leisure, finance, and corporate facilities calculated using POIs data, respectively. i is the coefficient factor of each feature, and the multi-agent feature importance results of each feature are used.
[0116] (2) Infrastructure factors
[0117] Medical POIs are used to describe the distribution of medical resources, including: primary hospitals, secondary hospitals, tertiary hospitals, and private clinics;
[0118] The distribution of nearby infrastructure is measured by calculating the number of life service POIs within 1500m of the cell, i.e., the point density.
[0119] Construct infrastructure factor evaluation indicators. The evaluation formula is:
[0120] F2(x, y)=I6Dis6(x, y)+I7Den7(x, y) Formula 4
[0121] Among them, F2(x, y) is the infrastructure factor evaluation index, Dis6 and Den7 are the distance to medical POIs and the density of life service POIs calculated using POIs data, respectively. i is the coefficient factor of each feature.
[0122] (3) Traffic factors
[0123] We used transportation facilities as keywords to capture POIs, using subway stations and bus stations as intra-city transportation points, and airports, railway stations, ports, and long-distance bus stations as inter-city transportation points to analyze the impact of different transportation facilities on housing prices.
[0124] The road network structure and road network-level data reflect the city's transportation network from a physical perspective and are closely related to the distribution of housing prices within the city. The road network data in this application includes national highways, provincial highways, expressways, urban main roads, urban secondary roads, urban branch roads, miscellaneous roads, rural roads, and county roads. By calculating the distance to each level of road, the road network structure near the community is also measured. By integrating these multiple factors, a traffic factor evaluation index is constructed. The evaluation formula is as follows:
[0125]
[0126] Among them, F3(x, y) is the traffic factor evaluation index, Dis8 is the distance to each transportation facility calculated using POIs data, and Dis i , i = 9, ..., 17 is the distance to national roads, provincial roads, expressways, urban main roads, urban secondary roads, urban branch roads, miscellaneous roads, rural roads, and county roads. i is the coefficient factor of each feature.
[0127] (4) Location factor
[0128] The location factor is a price-influencing factor of a residential area's location in the city. By capturing POIs of central government agencies, public security, procuratorial and judicial agencies, and government agencies at all levels, and calculating their point density as one of the location conditions for measuring a residential area, we constructed a location factor evaluation index. The evaluation formula is as follows:
[0129] F4(x, y)=1 18 Dis 18 (x, y) + I 19 Den 19 (x, y) Formula 6
[0130] Among them, F2(x, y) is the location factor evaluation index, Dis 18 、Den 19 They are the distance to the CBD of the convention and exhibition center and the density of government POIs, I i is the coefficient factor of each feature.
[0131] (5) Educational factors
[0132] By collecting the public high school enrollment rates of each school and those published on the education website as a quantitative indicator of the school's teaching staff, in terms of school district housing, by collecting the admission rules of each middle school provided by the government online, excluding schools that serve the entire region, the enrollment districts divided by each school were recorded, and the corresponding school district location information was obtained using the location query service provided by the map to construct an education factor evaluation index. The evaluation formula is as follows:
[0133] F5(x, y)=1 20 P 20 (x, y) Formula 7
[0134] Among them, F5(x, y) is the evaluation index of education factor, P 20 (x, y) is the enrollment rate of junior high schools in the school district where the community is located, I i is the coefficient factor of each feature. Figure 2 This is a chart showing the calculation results of educational factors and educational data for a certain city.
[0135] (6) Consumption level factor
[0136] Based on the POI data of Dianping.com, the average per capita consumption of restaurants near the community is obtained by applying the inverse distance weighted interpolation method to the food consumption point data. This is used to measure the consumption level of the community population and to construct a consumption level factor evaluation index. The evaluation formula is as follows:
[0137] F6(x, y)=1 21 P 21 (x, y) Formula 8
[0138] Among them, F6(x, y) is the consumption level factor evaluation index, P 21 (x, y) is the average per capita consumption of restaurants near the community, I i is the coefficient factor of each feature.
[0139] (7) Environmental factors
[0140] Based on Landsat8 data, the Normalized Difference Vegetation Index (NDVI) and the Modified Normalized Difference Water Index (MNDWI) were calculated. The average NDVI and MNDWI of each plot were calculated to obtain the vegetation and water cover of the community. The calculation formulas for NDVI and MNDWI are as follows:
[0141]
[0142] Among them, MIR is the mid-infrared band, NIR is the near-infrared band, R is the red band, and G is the green band;
[0143] Construct environmental factor evaluation indicators, and the evaluation formula is as follows:
[0144] F7(x, y)=1 22 NDVI+I 23 MNDWI formula 10
[0145] Among them, F7(x, y) is the environmental factor evaluation index, NDVI and MNDWI are the normalized vegetation index and improved normalized water index, I i is the coefficient factor of each feature. Figure 3 This is a graph showing the corrected normalized water index calculation results for a certain city.
[0146] (2) Quantification of the importance of impact factors based on multi-agents
[0147] Housing prices have a nonlinear functional relationship with location and environmental factors. We use feature importance evaluation in a multi-agent model to analyze the importance of each factor to housing prices in a nonlinear model. The multi-agent model is an integrated model in which the agent is a tree-like prediction model. The training sample dataset is input into the agent, and the sample is segmented by selecting some features. In each segmentation, a feature is selected for critical value division. Finally, samples at different leaf nodes have different category attributes, and the agent structure is adjusted through pruning.
[0148] When the number of features is large, splitting attributes and pruning cannot solve the problems of tree imbalance and overfitting. Multi-agent solves the problems of overfitting and local convergence by randomly selecting training sample sets and randomly selecting features for generating agents.
[0149] In the feature selection of generating agents, input variables are randomly selected, and part of the features of the input sample set are used to construct the agent. The selection process is also sampling with replacement, and a tree-structured data partition is constructed. By constructing agents multiple times, a multi-agent model is obtained. Through multiple random processes, the multi-agent model can better analyze the distribution of the overall data than a single agent. A multi-agent is a system composed of a series of agents h k (x)(k=1,…,N), the marginal function is defined as:
[0150] mg(X, Y)=av k (I(h k (X)=Y))-max j≠Y (av k (I(h k (X)=j))) Formula 11
[0151] Where I(·) is the indicator function, Y is the correct classification vector, j is the incorrect classification vector, and av k(·) means taking the average, and the marginal function indicates the degree to which the number of votes for the correct classification exceeds the maximum number of votes for the incorrect classification. The larger the value, the higher the confidence in the classifier;
[0152] The generalization error is used to show how well the model fits the training data, to prevent overfitting caused by the inability of the training data to estimate the distribution of the entire data and the model being overtrained to fit the training data. It is defined as:
[0153] PE=P X,Y (mg(X, Y)<0) Formula 12
[0154] Among them, P X,Y (·) indicates that the generalization error is obtained under the distribution of random variables X and Y;
[0155] The generalization error PE will tend to an upper bound, and the multi-agent algorithm will not suffer from serious overfitting problems as the number of agents increases. One of the known upper bounds of the generalization error is:
[0156]
[0157] Where avp is the average correlation between agents, s is the average strength of agents. To achieve good generalization performance of multi-agents, the correlation between agents should be reduced and the average strength of agents should be increased.
[0158] The grid search method is used to search, and the feature importance of the feature set is calculated after the optimal parameters are obtained to evaluate the impact of all features on the price of second-hand houses in the city. The feature importance is calculated multiple times and the ones with out-of-bag R-square less than 0 are removed. The 10 models are weighted by their out-of-bag R-square to obtain the average feature importance ( Figure 4 ).
[0159] The most important factors are education factor, transportation factor, consumption level, and location factor, followed by commercial development, infrastructure and environmental factors. The primary factor in considering buying a house has gradually become a good school district house, which is increasingly important for increasing housing prices. Secondly, compared with the commercial development factor, being close to transportation facilities and roads improves transportation convenience and accessibility, meeting daily work and shopping needs. In addition, the consumption level shows that communities with high-income groups have higher housing prices from the perspective of residents' income level. Secondly, it also shows the level of commercial development. Compared with the commercial development factor measured in quantity, points with high consumption levels are mostly high-end consumption areas, which means that the commercial level is higher than the commercial level. The number of points has a more significant impact on housing prices. The concentration of low-end small supermarkets and restaurants in urban villages also leads to a high commercial development factor in these areas, which can lead to misjudgments. Consumption level indirectly supplements the commercial development situation in the city. As the fourth most important location factor, it illustrates the importance of location and indirectly demonstrates the high concentration of urban development. The rich resources and convenient transportation in large cities make commercial agglomeration factors, infrastructure factors, and environmental factors relatively unimportant. People's daily lives are mainly about study and work, and they travel to meet their shopping and entertainment needs during holidays. Therefore, the commercial concentration, green space, and infrastructure near the community become relatively minor factors.
[0160] From the details, among the commercial development factors, the most important factors are leisure and entertainment, finance, corporate enterprises, food and shopping, which shows that the concentration of leisure and entertainment best reflects the commercial development. Residential areas near more leisure and entertainment areas have higher housing prices, followed by finance. The concentration area of banking and financial POIs is the CBD high-end commercial area, and the housing prices nearby are relatively high. Similarly, among the infrastructure factors, the concentration of life services is more important than the distance to the hospital. The concentration of life service facilities near the community will increase housing prices. In contrast, the travel convenience near the hospital decreases due to the large number of vehicles entering and leaving, and the noise generated by ambulance sirens has a negative impact on people's living standards to a certain extent, thereby affecting people's demand for housing. Among the traffic factors, the distance to traffic facilities is relatively close to the distance to roads at all levels. Distance is less important, which also reflects that people with lower income levels tend to choose public transportation, and the proximity to subway stations and bus stops becomes an important factor they consider, while people with higher income levels tend to choose taxis or drive, and road accessibility becomes a factor they consider more. On the other hand, the city has a highly developed public transportation system, and the convenience of using public transportation between different communities is not much different. The distance to transportation facilities has a smaller impact on housing prices. Among location factors, the distance to the CBD is more important than the density of government POIs. The highly concentrated and monocentric nature of urban housing prices means that housing prices basically decrease with increasing distance from the city center. Although each district may have local extreme values, the highest values in the entire district are basically concentrated in specific areas and there are huge differences between them and other districts.
[0161] 3. Establishing a model for the spatial distribution of urban second-hand housing prices
[0162] (1) Hedonic price regression model for urban second-hand housing
[0163] Hedonic price regression is used to explore the relationship between the built environment and second-hand housing prices. After model training, the density characteristics of POIs can be used to calculate housing price information. Nonlinear functional relationships are fitted by minimizing empirical risk.
[0164] There is a training sample set (X,Y) = {x 1i ,…,x si ,y i}, i=1,…,N, where N is the number of samples, X={x 1i ,…,x si} is the sample feature set, s is the number of features, y i For the corresponding output value, the relationship is:
[0165]
[0166] in is a kernel function that maps the X vector to a high-dimensional space and makes it have a linear relationship with the output value Y: the ω and b parameters are solved using the following formula:
[0167]
[0168] where ξ i and is a relaxation factor that reduces the noise in the data to avoid overfitting; ∈ is the bound of the loss function, which allows a reasonable error between the regression result and the true value; C is a parameter that balances the regression result and the relaxation variable;
[0169] Solve the constrained quadratic programming problem and find ω and b as follows:
[0170]
[0171] where α i , is the minimization solution of J(ω, e), and finally the model expression is obtained:
[0172]
[0173] in is the kernel function.
[0174] (2) Geographically Weighted Regression Model for Second-Hand Housing Prices
[0175] By inputting the coordinates of the variables into the model as parameters and using non-parametric estimation to obtain the corresponding local function estimate at each location, we can finely analyze the distribution of the dependent variable. Based on the first law of geography, we model different regions separately to reflect the spatial heterogeneity of the distribution of housing prices and influencing factors, as well as the spatial heterogeneity of the relationship between housing prices and influencing factors. The model is expressed as:
[0176]
[0177] Where, (u i , v i ) is the geographical coordinate of the center point of the i-th road network segmentation object, x ik is the characteristic variable of the i-th object, β0(u i , v i ) is the sum of the constants of other influencing factors of the characteristic variable, β k (u i , v i ) is the feature x ik The regression coefficient, ε i is a random error term; a linear function is used to characterize the characteristic regression coefficient of housing prices, and the expression is:
[0178] β t=λ t0 +λ t1 u i +λ t2 v i Formula 21
[0179] For any point in the region (u i , v i ) is estimated by least squares:
[0180]
[0181] Where X is the impact factor matrix, W(u i , v i ) is a weight matrix, which is composed of the monotonically decreasing function value of the spatial position distance between the regression object and other surrounding objects. The kernel function expression is as follows:
[0182]
[0183] Where d is the Euclidean distance between objects and h is the optimal bandwidth, which is obtained by the cross determination method:
[0184]
[0185] in is the value of the house obtained after fitting the i-th object, y i is the true value of the house price at object i.
[0186] (3) Calculation and analysis of the spatial distribution of housing prices
[0187] Data preprocessing was performed using ArcGIS 10.0. The multi-agent algorithm and hedonic price regression algorithm were implemented using the scikit-learn library in Python. The hedonic price regression model for urban second-hand housing was implemented using SPSS. The geographically weighted regression model for second-hand housing prices was implemented using the Geographically Weighted Regression tool in the Spatial Relationship Modeling toolset in the Spatial Computing Toolbox in ArcMap.
[0188] Normalize the feature data by subtracting the sample mean from the sample data and dividing the difference by the sample standard deviation, so that the sample mean becomes 0 and the standard deviation becomes 1, thus removing the dimension. The formula is as follows:
[0189]
[0190] Where μ is the mean of the sample data, σ is the standard deviation of the sample data;
[0191] The Root Mean Square Error (RMSE) and R-squared are used as accuracy evaluation indicators. The Root Mean Square Error represents the difference between the predicted result and the actual result. It is related to the dimension of the regression result and the closer it is to 0, the better. The value range of R-squared is between [0, 1]. The closer it is to 1, the better the model. The formula is as follows:
[0192] The model is validated by cross-validation, which divides the sample set into two parts, and uses only one part of the sample data to train the model parameters, while the other part of the data is used to verify the generalization ability of the model. The k-fold cross-validation method is used, which divides the sample into k parts each time, selects one of them as the validation sample, and the others are untrained samples. The average RMSE and R-squared are calculated as the accuracy evaluation results.
[0193] Using the housing price influencing factor system constructed in this application as the independent variable, regression prediction is performed using the multi-agent regression model, the characteristic price regression model, the linear regression model, and the geographically weighted regression model to obtain the results of urban second-hand housing prices in the entire city. After evaluating the model.
[0194] Hedonic price regression has different effects when using different kernel functions. The Gaussian kernel function for hedonic price regression, suitable for nonlinear relationships, has an RMSE of 10651.342 and an R² of 0.362, while the linear kernel function for hedonic price regression, suitable for linear relationships, has an RMSE of 10979.052 and an R² of 0.346. The accuracy of the multi-agent model is RMSE of 12888.228 and an R² of 0.278. In comparison, the accuracy of the linear regression model is RMSE of 17502.941 and an R² of 0.296. Comparing the Gaussian kernel SVM with the linear kernel SVM, and the SVM, RFA nonlinear fitting model, and linear regression, shows that the relationship between influencing factors and housing prices is more consistent with a nonlinear relationship than a linear relationship, indicating the complex relationship between transportation, education, consumption level, and housing prices.
[0195] Based on the spatial heterogeneity of urban housing prices, a geographically weighted regression model was introduced for comparison. This model models each spatial object and introduces the results of this object and other objects as parameters into the model to obtain a model that is more in line with geographical laws. The results are as follows: RMSE -15748.661, R^2 = 0.446. This model is more accurate than the best linear regression model of zoning modeling, indicating that there are differences in the relationship between housing prices and their influencing factors in different geographical locations.
[0196] To further analyze the spatial clustering and differentiation of urban housing prices, we calculated the local Moran's I index and combined it with hypothesis testing to identify high-value clusters, low-value clusters, and outlier areas for pre-owned housing. This analysis was implemented using the Cluster and Profile Analysis module in the ArcGIS Spatial Computing Toolbox. The following results were obtained by analyzing the results of the most accurate zonal RFA model. HH represents high values clustered near high values, LL represents low values clustered near low values, L represents low values clustered near high values, LH represents high values clustered near low values, and Nu11 represents random distribution of housing prices.
[0197] The spatial distribution of housing prices not only shows continuity in accordance with the first law of geology, but also shows discontinuity in spatial distribution due to the asynchrony of the urbanization process and various historical reasons.
[0198] This application constructs a hedonic price model for urban second-hand housing based on multiple micro-factors. Combining web crawlers and machine learning methods, this paper analyzes the distribution of urban housing prices and the relationship between housing prices and influencing factors from various perspectives, drawing the following conclusions:
[0199] (1) Through the calculation and analysis of urban housing price data, it is found that the mean values of housing prices in different districts are different. Further, through the analysis of the cumulative histogram and normal QQ plot of housing prices in the entire city, it is found that the distribution of urban housing prices does not conform to the normal distribution, and there are more high values. After the original data are transformed by the log function, the data conform to the normal distribution. At the same time, through the spatial autocorrelation analysis, it is found that there is a spatial clustering phenomenon in urban housing prices, with a continuous surface fitted by the function rather than a random distribution, which shows the feasibility of modeling. Through the spatial heterogeneity analysis, it is found that there is a certain degree of difference in housing prices in various districts of the city.
[0200] (2) Since this application is to analyze the fine distribution of urban housing prices, relatively micro-influencing factors are used to explore the relationship between housing prices and influencing factors during the feature construction process, including seven aspects: commercial development, transportation, infrastructure, location, education, environment and residents' consumption level. Data acquisition is combined with web crawlers, and spatial analysis tools are used to quantify the influencing factors. The multi-agent feature importance model is further used to analyze the degree of influence of each influencing factor on housing prices in a nonlinear model. It is found that the most important factors are education factors, transportation factors, consumption level, location factors, followed by commercial development, infrastructure and environmental factors. The degree of influence of the internal structure of each factor on housing prices is analyzed in detail, explaining the relationship between urban housing prices and influencing factors.
[0201] (3) A nonlinear characteristic price model based on network big data is proposed to obtain the fine spatial distribution of urban second-hand housing prices. At the same time, different nonlinear models, linear models, and geographically weighted regression models are compared, and a nonlinear characteristic price model based on partitioning is proposed to obtain the distribution patterns of housing prices in different parts of the city. After partitioning modeling, the accuracy of each algorithm is improved, and the multi-agent model has the greatest improvement. At the same time, a more accurate result of the fine distribution of urban housing prices is also obtained.
[0202] (4) Finally, the spatial clustering of urban housing prices is analyzed based on the results of a high-precision partitioned multi-agent model. Local Moran's I is used to analyze the high-value and low-value areas of urban housing prices, thereby analyzing the spatial pattern of urban housing prices. Furthermore, by analyzing the distribution of outlier areas, recommendations are provided for urban renewal and for the construction of high-end residential communities by real estate developers.
Claims
1. A multi-agent network big data urban housing price analysis and calculation method, characterized by: First, we analyze the multi-factor spatial data of urban second-hand housing prices, using housing price spatial heterogeneity analysis and housing price spatial autocorrelation analysis. Then, we model the factors affecting urban second-hand housing prices, including constructing urban second-hand housing price influencing factors and quantifying the importance of influencing factors based on multi-agents. The construction of urban second-hand housing price influencing factors measures the locational environmental factors that affect housing prices from seven aspects: business development, infrastructure, transportation, location, education, consumption level, and environment. Finally, we establish a model for the spatial distribution of urban second-hand housing prices, including an urban second-hand housing hedonic price regression model and a second-hand housing geographically weighted regression model, and jointly calculate and analyze the spatial distribution of housing prices. 1) Analyze multi-factor spatial data on urban second-hand housing prices: Analyze the original housing price data based on zoning and histogram calculations to obtain the basic spatial distribution and calculation characteristics of urban second-hand housing prices; Using geographic detectors, we conducted a spatial heterogeneity analysis of second-hand housing prices in various administrative districts of the city, obtaining the distribution differences of housing prices between different administrative districts. We also conducted a global spatial autocorrelation analysis of second-hand housing prices in the city using the Moran's I tool, obtaining the spatial autocorrelation and spatial clustering characteristics of second-hand housing prices in the city. 2) Modeling factors influencing urban second-hand housing prices: Build a model for the impact of locational environmental factors on housing prices. Utilize POI data, review POI data, road network data, and Landsat data to quantify the distribution of various social, economic, and locational factors near residential areas, including commercial development, transportation, infrastructure, location, education, environment, and resident consumption levels. These characteristics are then integrated into a multi-intelligence assessment model to determine the contribution of various characteristics to the spatial variation of housing prices in a nonlinear model. 3) Spatial distribution model of urban second-hand housing prices: Based on the results of multi-factor spatial analysis and influencing factor analysis, different characteristic price models are established to obtain high-precision urban housing price distribution results. Among them, the partitioned nonlinear characteristic price model has higher accuracy than the linear model and the global modeling model. The local Moran's I is further used to analyze the spatial layout of the urban housing price market and distinguish high-value areas, transition areas, and low-value areas.
2. The multi-agent network big data city housing price analysis and calculation method according to claim 1 is characterized in that: Analysis of spatial heterogeneity of housing prices: Based on geographic detectors, the differences in housing prices in different administrative regions are calculated. Under the condition of discrete independent variables, the dependent variable is divided into blocks. The differences between different blocks are spatial heterogeneity. The q value of the differentiation detector detects the spatial heterogeneity of the dependent variable. The expression is: Where L is the stratification of factor X, N h and N are the number of samples in each layer and the whole area, and σ 2 is the variance of the dependent variable of each district and the whole district, and the value range of q is [0, 1]. The larger the value, the greater the difference of the dependent variable among the districts. Analysis of spatial autocorrelation of housing prices: Based on the spatial clustering of housing price datasets, the spatial autocorrelation is calculated, which is expressed as follows: where z i is the difference between the attribute value of sample i and the sample mean, w i,j is the spatial weight between sample i and sample j, n is the number of samples, and when the value of I is 0, the housing price has nothing to do with its spatial position, and the housing price is randomly distributed in space; when the value is closer to 1, it indicates that there are high values and high value clusters of housing prices in space. On the contrary, the closer it is to -1, it indicates that there is a step-like fault in the distribution of housing prices, and there are abnormal land price areas around high-priced areas.
3. The multi-agent network big data city housing price analysis and calculation method according to claim 1 is characterized in that: Infrastructure factor: Medical POIs are used to describe the distribution of medical resources, including primary hospitals, secondary hospitals, tertiary hospitals, and private clinics. The distribution of nearby infrastructure is measured by calculating the number of life service POIs within 1500m of the cell, i.e., the point density. Construct infrastructure factor evaluation indicators. The evaluation formula is: F2(x, y)=I6Dis6(x, y)+I7Den7(x, y) Formula 4 Among them, F2(x, y) is the infrastructure factor evaluation index, Dis6 and Den7 are the distance to medical POIs and the density of life service POIs calculated using POIs data, respectively. i is the coefficient factor of each feature.
4. The multi-agent network big data city housing price analysis and calculation method according to claim 1 is characterized in that: Transportation factors: We use transportation facilities as keywords to capture POIs. We use subway stations and bus stations as intra-city transportation points, and airports, railway stations, ports, and long-distance bus stations as inter-city transportation points to analyze the impact of different transportation facilities on housing prices. The road network structure and road network-level data reflect the city's transportation network from a physical perspective and are closely related to the distribution of housing prices within the city. The road network data in this application includes national highways, provincial highways, expressways, urban main roads, urban secondary roads, urban branch roads, miscellaneous roads, rural roads, and county roads. By calculating the distance to each level of road, the road network structure near the community is also measured. By integrating these multiple factors, a traffic factor evaluation index is constructed. The evaluation formula is as follows: Among them, F3(x, y) is the traffic factor evaluation index, Dis8 is the distance to each transportation facility calculated using POIs data, and Dis i , i = 9, ..., 17 is the distance to national roads, provincial roads, expressways, urban main roads, urban secondary roads, urban branch roads, miscellaneous roads, rural roads, and county roads. i is the coefficient factor of each feature.
5. The multi-agent network big data city housing price analysis and calculation method according to claim 1 is characterized in that: Location factor: The location factor is a price-influencing factor of a residential area's location within the city. By capturing POIs of central government agencies, public security, procuratorial, and judicial institutions, and government agencies at all levels, and calculating their point density as one of the location conditions for measuring a residential area, we construct a location factor evaluation index. The evaluation formula is as follows: F4(x, y) = I 18 Dis 18 (x, y) + I 19 Den 19 (x, y) Equation 6 Among them, F2(x, y) is the location factor evaluation index, Dis 18 、Den 19 They are the distance to the CBD of the convention and exhibition center and the density of government POIs, I i is the coefficient factor of each feature.
6. The multi-agent network big data city housing price analysis and calculation method according to claim 1 is characterized in that: Education factor: By collecting the public high school enrollment rates of each school and the schools published on the education website, we use them as a quantitative indicator of the school's teaching staff. In terms of school district housing, we collect the admission rules of each middle school provided by the government online. Excluding schools that serve the entire region, we record the enrollment districts divided by each school. We use the location query service provided by the map to obtain the corresponding school district location information to construct an education factor evaluation index. The evaluation formula is as follows: F5(x, y) = I 20 P 20 (x, y) Equation 7 Among them, F5(x, y) is the evaluation index of education factor, P 20 (x, y) is the enrollment rate of junior high schools in the school district where the community is located, I i is the coefficient factor of each feature.
7. The multi-agent network big data city housing price analysis and calculation method according to claim 1 is characterized in that: Consumption level factor: Based on the POI data of Dianping.com, the average per capita consumption of restaurants near the community is obtained by applying the inverse distance weighted interpolation method to the food consumption point data. This is used to measure the consumption level of the community population and construct a consumption level factor evaluation index. The evaluation formula is as follows: F6(x, y) = I 21 P 21 (x, y) Equation 8 Among them, F6(x, y) is the consumption level factor evaluation index, P 21 (x, y) is the average per capita consumption of restaurants near the community, I i is the coefficient factor of each feature. (7) Environmental factors Based on Landsat8 data, the Normalized Difference Vegetation Index (NDVI) and the Modified Normalized Difference Water Index (MNDWI) were calculated. The average NDVI and MNDWI of each plot were calculated to obtain the vegetation and water cover of the community. The calculation formulas for NDVI and MNDWI are as follows: Among them, MIR is the mid-infrared band, NIR is the near-infrared band, R is the red band, and G is the green band; Construct environmental factor evaluation indicators, and the evaluation formula is as follows: F7(x, y) = I 22 NDVI + I 23 MNDWI Equation 10 Among them, F7(x, y) is the environmental factor evaluation index, NDVI and MNDWI are the normalized vegetation index and improved normalized water index, I i is the coefficient factor of each feature.
8. The multi-agent network big data city housing price analysis and calculation method according to claim 1 is characterized in that: Quantifying the importance of influencing factors based on multi-agents: Housing prices have a nonlinear functional relationship with location and environmental factors. The importance of each factor in the nonlinear model of housing prices is analyzed using feature importance evaluation in multi-agents. The multi-agent model is an integrated model in which the agent is a tree-like prediction model. The training sample dataset is input into the agent, and the sample is segmented by selecting some features. In each segmentation, a feature is selected for critical value division. Finally, samples at different leaf nodes have different category attributes, and the agent structure is adjusted through pruning. In the feature selection of generating agents, input variables are randomly selected, and part of the features of the input sample set are used to construct the agent. The selection process is also sampling with replacement, and a tree-structured data partition is constructed. By constructing agents multiple times, a multi-agent model is obtained. Through multiple random processes, the multi-agent model can better analyze the distribution of the overall data than a single agent. A multi-agent is a system composed of a series of agents h k (x)(k=1,…,N), the marginal function is defined as: mg(X, Y) = av k (I(hk(X) = Y)) - max j≠Y (av k (I(h k (X) = j))) Equation 11 Where I(·) is the indicator function, Y is the correct classification vector, j is the incorrect classification vector, and av k (·) means taking the average, and the marginal function indicates the degree to which the number of votes for the correct classification exceeds the maximum number of votes for the incorrect classification. The larger the value, the higher the confidence in the classifier; The generalization error is used to show how well the model fits the training data, to prevent overfitting caused by the inability of the training data to estimate the distribution of the entire data and the model being overtrained to fit the training data. It is defined as: PE = P X,Y (mg(X, Y) < 0) Equation 12 Among them, P X,Y (·) indicates that the generalization error is obtained under the distribution of random variables X and Y; The generalization error PE will tend to an upper bound, and the multi-agent algorithm will not suffer from serious overfitting problems as the number of agents increases. One of the known upper bounds of the generalization error is: Where avp is the average correlation between agents, s is the average strength of agents. To achieve good generalization performance of multi-agents, the correlation between agents should be reduced and the average strength of agents should be increased. The grid search method is used for search. After obtaining the optimal parameters, the feature importance of the feature set is calculated to evaluate the impact of all features on the price of second-hand houses in the city. The feature importance is calculated multiple times and the ones with out-of-bag R-square less than 0 are removed. The 10 models are weighted according to their out-of-bag R-square to obtain the average feature importance.
9. The multi-agent network big data city housing price analysis and calculation method according to claim 1 is characterized in that: Urban Second-Hand Housing Hedonic Price Regression Model: This model uses hedonic price regression to explore the relationship between the built environment and second-hand housing prices. After model training, it can use POI density characteristics to calculate housing price information. It also fits nonlinear functional relationships by minimizing empirical risk. There is a training sample set (X, Y) = {x 1i ,…,x si ,y i }, i = 1, ..., N, where N is the number of samples, X = {x 1i ,…,x si } is the sample feature set, s is the number of features, y i For the corresponding output value, the relationship is: in is a kernel function that maps the X vector to a high-dimensional space and makes it have a linear relationship with the output value Y: the ω and b parameters are solved using the following formula: where ξ i and is a relaxation factor that reduces the noise in the data to avoid overfitting; ∈ is the bound of the loss function, which allows a reasonable error between the regression result and the true value; C is a parameter that balances the regression result and the relaxation variable; Solve the constrained quadratic programming problem and find ω and b as follows: where α i , is the minimization solution of J(ω, e), and finally the model expression is obtained: in is the kernel function.
10. The multi-agent network big data city housing price analysis and calculation method according to claim 1 is characterized in that: Geographically weighted regression model for second-hand housing prices: By inputting the coordinates of the variables into the model as parameters and using non-parametric estimation to obtain the corresponding local function estimate at each location, the distribution of the dependent variable is analyzed in detail. Based on the first law of geography, different regions are modeled separately to reflect the spatial heterogeneity of the distribution of housing prices and influencing factors, as well as the spatial heterogeneity of the relationship between housing prices and influencing factors. The model is expressed as: Where, (u i , v i ) is the geographical coordinate of the center point of the i-th road network segmentation object, x ik is the characteristic variable of the i-th object, β0(u i , v i ) is the sum of the constants of other influencing factors of the characteristic variables, β k (u i , v i ) is the feature x ik The regression coefficient, ε i is a random error term; a linear function is used to characterize the characteristic regression coefficient of housing prices, and the expression is: b t =λ t0 +λ t1 you i +λ t2 v i formula 21 For any point in the region (u i , v i ) is estimated by least squares: Where X is the impact factor matrix, W(u i , v i ) is a weight matrix, which is composed of the monotonically decreasing function value of the spatial position distance between the regression object and other surrounding objects. The kernel function expression is as follows: Where d is the Euclidean distance between objects, and h is the optimal bandwidth, which is obtained by the cross determination method: in is the value of the house obtained after fitting the i-th object, y i is the true value of the house price at object i.