House data processing method and device, medium and product
By dividing data into community and housing unit dimensions, and combining decision trees and pre-built value estimation models, the value of houses is accurately determined, which solves the problem of insufficient accuracy in existing methods for determining the unit price of houses, and achieves high precision and consistency in house value assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA CONSTRUCTION BANK
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, the accuracy of methods for determining the unit price of a house is poor, resulting in a large discrepancy between the determined house value and the actual transaction price, making it difficult to accurately reflect the actual value of a single property.
By dividing data samples into community-level and housing-level categories, we conduct in-depth analysis of the relationship between various marginal influencing factors and housing prices, determine the marginal factor coefficients of each marginal influencing factor, and combine them with core influencing factors to accurately calculate housing values using decision trees and pre-built value estimation models.
It significantly improves the accuracy of housing valuation, breaks through the limitations of traditional reliance on static expert rules, realizes the evolution from unified standards to community-level dynamic modeling, and improves the consistency between housing valuation results and actual transaction prices.
Smart Images

Figure CN121836767A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, device, medium and product for processing housing data. Background Technology
[0002] In the real estate transaction sector, the benchmark price of a property (i.e., the average price of similar properties in the same area) is often insufficient to directly and accurately reflect the actual value of a single property. Determining the value of a single property requires a comprehensive analysis of various influencing factors.
[0003] In existing technologies, the average price of a property is usually used as the base price, and then coefficients of influencing factors such as area type, orientation, and decoration are combined to calculate the unit price of the house. Among them, the coefficients of these influencing factors are mainly determined based on unified standards formulated by experts.
[0004] However, the existing methods for determining the unit price of houses are not very accurate, resulting in a large discrepancy between the determined house value and the actual transaction price. Summary of the Invention
[0005] This application provides a housing data processing method, device, medium, and product, which uses the method of dividing data samples by community dimension and housing source dimension to determine the edge factor coefficients of each edge influencing factor, thereby improving the accuracy of the determined housing value and solving the problem of large deviation between the housing value determined by the existing housing unit price determination method and the actual transaction price.
[0006] In a first aspect, embodiments of this application provide a method for processing housing data, the method comprising:
[0007] Obtain a sample dataset of houses in the area where the target house is located;
[0008] Based on the housing sample dataset, determine the sample field information that meets the preset coverage conditions, including various housing value influencing factors;
[0009] By analyzing the sample field information using decision trees, the core influencing factors and multiple different marginal influencing factors are identified.
[0010] Based on the housing sample dataset, housing sample data corresponding to communities where the value difference reaches a preset difference threshold are determined as training sample datasets.
[0011] The training sample dataset is binned in two dimensions to obtain data samples divided by the community dimension and data samples divided by the housing dimension.
[0012] Statistical analysis was performed on data samples divided by community dimension and data samples divided by housing dimension to obtain the marginal factor coefficients of each marginal influence factor.
[0013] Based on the data sample division by community dimension, data sample division by housing dimension, and the pre-built value estimation model, the model estimate of the unit price of the community sample containing core influencing factors is determined.
[0014] The target house value is determined by estimating the sample unit price of the community and the marginal factor coefficients of each marginal influencing factor based on the model.
[0015] In one possible implementation, the sample field information is analyzed using a decision tree to determine the core influencing factors and multiple distinct marginal influencing factors, including:
[0016] By using a decision tree and pre-set analysis tools on the sample field information, we obtained various factors affecting house value and the corresponding impact scores of each factor.
[0017] The housing value influencing factors and their corresponding influence scores are compared numerically to identify the housing value influencing factor with the largest influence score as the core influencing factor, and the other housing value influencing factors besides the core influencing factor are identified as marginal influencing factors.
[0018] In one possible implementation, the training sample dataset is binned in two dimensions to obtain data samples divided by neighborhood dimension and data samples divided by housing unit dimension, including:
[0019] The training sample dataset is divided into cell-level data samples by the median area, resulting in a first-area cell dataset and a second-area cell dataset, where the first area is larger than the second area.
[0020] The training sample dataset is binned by house area in the house source dimension to obtain multiple sets of house sample data with different house types as house source dimension partitioning data samples.
[0021] In one possible implementation, statistical analysis is performed on the data samples segmented by community dimension and the data samples segmented by housing unit dimension to obtain the marginal factor coefficients of each marginal influence factor, including:
[0022] Based on the data samples divided by the community dimension, a correlation analysis of housing area and unit price was conducted to obtain the results of the correlation analysis between housing area and price at the community dimension.
[0023] Based on the correlation analysis results between housing area and price at the community level and the data samples divided at the housing source level, statistical calculations were performed to obtain the area coefficient, decoration coefficient, orientation coefficient, building type coefficient, unit type coefficient, and floor coefficient as the marginal factor coefficients of each marginal influencing factor.
[0024] In one possible implementation, based on data sample segmentation at the community level, data sample segmentation at the housing level, and a pre-built value estimation model, the model-estimated unit price of the community sample, including core influencing factors, is determined, including:
[0025] Based on the data sample division by community dimension and the data sample division by housing dimension, the average area and average unit price within the first fixed time period are determined.
[0026] The benchmark price for the second fixed time period is determined by dividing the data samples according to the community dimension and the housing unit dimension.
[0027] Data samples are divided into two categories based on the community dimension and the housing dimension to determine the binned composite data.
[0028] Data samples are divided into two categories based on the community dimension and the housing unit dimension to determine the sample area;
[0029] Based on the average area and average unit price in the first fixed time period, the benchmark price in the second fixed time period, the composite data of different buckets, the sample area, and the pre-built value estimation model, the model estimate of the sample unit price of the community containing the core influencing factors is determined.
[0030] In one possible implementation, the formula for determining the target house value is as follows: Based on the model-estimated unit price of the neighborhood sample and the marginal factor coefficients of each marginal influence factor.
[0031] Y = Model estimated unit price of the community sample × area coefficient × decoration coefficient × orientation coefficient × building type coefficient × unit type coefficient × floor coefficient.
[0032] In one possible implementation, before determining the sample field information that meets the preset coverage conditions based on the housing sample dataset, the method further includes:
[0033] The house sample dataset is denoised to obtain an interference-resistant house sample dataset.
[0034] Secondly, embodiments of this application provide a housing data processing apparatus, the apparatus comprising:
[0035] The acquisition module is used to acquire a sample dataset of houses in the area where the target house is located.
[0036] The impact factor determination module is used to determine the sample field information that meets the preset coverage conditions based on the housing sample dataset. The sample field information includes a variety of housing value impact factors.
[0037] The impact factor determination module is also used to analyze sample field information through decision trees to determine core impact factors and multiple different marginal impact factors;
[0038] The training data determination module is used to determine the housing sample data corresponding to communities where the value difference reaches a preset difference threshold as the training sample dataset, based on the housing sample dataset.
[0039] The training data determination module is also used to perform two-dimensional bucketing on the training sample dataset to obtain data samples divided by community dimension and data samples divided by housing source dimension.
[0040] The coefficient determination module is used to perform statistical analysis on data samples divided by community dimension and data samples divided by housing dimension to obtain the edge factor coefficients of each edge influence factor;
[0041] The core unit price determination module is used to divide data samples according to the community dimension, the housing dimension, and the pre-built value estimation model to determine the model estimated unit price of the community sample, which includes core influencing factors.
[0042] The housing value determination module is used to determine the target housing value based on the estimated unit price of the sample housing area and the marginal factor coefficients of each marginal influencing factor according to the model.
[0043] In one possible implementation, the influence factor determination module is also used to obtain multiple house value influence factors and the influence scores corresponding to each house value influence factor by using a preset analysis tool on the sample field information through a decision tree.
[0044] The impact factor determination module is also used to perform numerical comparison processing based on each housing value impact factor and the corresponding impact score of each housing value impact factor, to obtain the housing value impact factor corresponding to the largest impact score as the core impact factor, and to determine the other housing value impact factors besides the core impact factor as marginal impact factors.
[0045] In one possible implementation, the training data determination module is further configured to perform area median partitioning on the training sample dataset in the cell dimension to obtain a first area cell dataset and a second area cell dataset as cell dimension partitioning data samples, wherein the first area is larger than the second area.
[0046] The training data determination module is also used to perform house area binning on the training sample dataset in the house source dimension, and obtain multiple sets of house sample data with different house source types as house source dimension partitioning data samples.
[0047] In one possible implementation, the coefficient determination module is also used to divide the data samples according to the community dimension to perform a correlation analysis of housing area and unit price, and obtain the correlation analysis results of housing area and price under the community dimension.
[0048] The coefficient determination module is also used to perform statistical calculations based on the correlation analysis results between housing area and price under the community dimension and the data samples divided by the housing source dimension, to obtain the area coefficient, decoration coefficient, orientation coefficient, building type coefficient, unit type coefficient and floor coefficient as the edge factor coefficients of each edge influencing factor.
[0049] In one possible implementation, the core unit price determination module is also used to divide the data samples according to the community dimension and the housing dimension to determine the average area and average unit price within a first fixed time period.
[0050] The core unit price determination module is also used to divide data samples according to the community dimension and the housing dimension to determine the benchmark price in the second fixed time period;
[0051] The core unit price determination module is also used to divide data samples according to the community dimension and the housing dimension to determine the binned composite data;
[0052] The core unit price determination module is also used to divide data samples according to the community dimension and the housing dimension to determine the sample area;
[0053] The core unit price determination module is also used to determine the model estimated unit price of the community sample, which includes core influencing factors, based on the average area and average unit price in the first fixed time period, the benchmark price in the second fixed time period, the binned composite data, the sample area, and the pre-built value estimation model.
[0054] In one possible implementation, in the housing value determination module, the calculation formula for determining the target housing value is as follows: based on the model-estimated sample unit price of the community and the marginal factor coefficients of each marginal influence factor:
[0055] Y = Model estimated unit price of the community sample × area coefficient × decoration coefficient × orientation coefficient × building type coefficient × unit type coefficient × floor coefficient.
[0056] In one possible implementation, the device further includes a preprocessing module;
[0057] The preprocessing module is used to denoise the house sample dataset to obtain an interference-resistant house sample dataset.
[0058] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0059] The memory stores the instructions that the computer executes;
[0060] The processor executes computer execution instructions stored in memory to implement a housing data processing method according to the first aspect of the invention.
[0061] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement a housing data processing method according to the first aspect of the invention.
[0062] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, is used to implement a housing data processing method according to the first aspect of the invention.
[0063] The housing data processing method, device, medium, and product provided in this application include: First, acquiring a housing sample dataset of the area where the target housing is located; then, determining sample field information that meets a preset coverage condition based on the housing sample dataset, wherein the sample field information includes multiple housing value influencing factors; then, analyzing the sample field information through a decision tree to determine core influencing factors and multiple different marginal influencing factors; and, based on the housing sample dataset, determining housing sample data corresponding to communities where the value difference reaches a preset difference threshold as training sample dataset; subsequently, performing two-dimensional binning on the training sample dataset to obtain community-level segmented data samples and housing-level segmented data samples; and performing statistical analysis on the community-level segmented data samples and housing-level segmented data samples to obtain the marginal factor coefficients of each marginal influencing factor; then, based on the community-level segmented data samples, housing-level segmented data samples, and a pre-built value estimation model, determining the model-estimated community sample unit price containing the core influencing factors; finally, determining the target housing value based on the model-estimated community sample unit price and the marginal factor coefficients of each marginal influencing factor. The following technical effects were achieved: Based on feature importance analysis, the area factor was identified as the core variable with the most significant impact on the unit price of houses. A two-dimensional segmentation mechanism was then introduced: at the community level, communities were divided into first-area communities and second-area communities based on their overall area distribution; at the housing unit level, individual housing units were further subdivided into five categories based on area. Finally, for different community types, data samples were segmented based on the community and housing unit dimensions, and a pre-built value estimation model was established. The estimated unit price of community samples was determined by removing interference from various marginal factors. Combined with the marginal factor coefficients of each marginal factor, the target house value was determined. This approach breaks through the limitations of traditional reliance on static expert rules, achieving an evolution from determining marginal factor coefficients using a unified standard to dynamic modeling at the community level, significantly improving the consistency between house valuation results and actual transaction prices. By comprehensively considering the impact of multiple influencing factors on house value, the accuracy of house valuation was effectively improved, providing strong support for real estate transactions and market supervision. Attached Figure Description
[0064] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0065] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0066] Figure 1 This is a schematic diagram illustrating an application scenario of a housing data processing method provided in an embodiment of this application.
[0067] Figure 2 A schematic flowchart illustrating a housing data processing method provided in an embodiment of this application;
[0068] Figure 3 This is a schematic diagram of the structure of a housing data processing device provided in an embodiment of this application;
[0069] Figure 4 This is a schematic diagram of the structure of an electronic device provided in this application.
[0070] Figure label:
[0071] 101-Terminal; 102-Server; 310-Acquisition Module; 320-Influence Factor Determination Module; 330-Training Data Determination Module; 340-Coefficient Determination Module; 350-Core Unit Price Determination Module; 360-House Value Determination Module; 410-Processor; 420-Memory; 430-Communication Components; 440-Bus. Detailed Implementation
[0072] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0073] In the embodiments of this application, the terms "first" and "second" are used to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply difference. It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate that something is being used as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being better or more advantageous than other embodiments or design schemes. Specifically, the use of "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner. In the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more.
[0074] It should be noted that the phrase "at...time" in the embodiments of this application can refer to the instant at which a certain situation occurs, or to a period of time after the occurrence of a certain situation; the embodiments of this application do not specifically limit this. Furthermore, the housing data processing method provided in the embodiments of this application is merely an example; a housing data processing method may include more or fewer elements.
[0075] The collection, storage, use, processing, transmission, provision, and disclosure of financial data or user data involved in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0076] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0077] In real estate transactions and valuation practices, the benchmark price of a property (i.e., the average transaction or listing price of similar properties in the same area) is usually used as the starting point for valuation. However, since this benchmark price is difficult to accurately reflect the individual value differences of a single property, existing technologies generally adopt a "one price per unit" adjustment model. This involves multiplying the average price of the property by several preset influencing factor coefficients (such as area type, orientation, decoration level, floor, and unit type) to adjust the price per unit.
[0078] Currently, the coefficients of the aforementioned influencing factors are mostly based on unified standards developed by industry experts and roughly calibrated using historical statistical values. For example, the area factor is simply divided into four levels from smallest to largest, without considering the structural differences in the relationship between area and price within different cities and communities. This ignores the true impact of the micro-environment of the property (such as community dimensions and property dimensions) on the price, leading to a significant deviation between the valuation results and actual transaction prices, especially in highly heterogeneous markets.
[0079] Therefore, this method of determining influencing factor coefficients based on a unified standard has obvious flaws. Because different communities and properties have their own unique characteristics and attributes, a unified standard cannot fully consider these differences, resulting in poor accuracy of the determined unit price. This leads to a significant discrepancy between the determined property value and the actual transaction price, failing to meet the demand for accurate valuation in real estate transactions.
[0080] Based on this, embodiments of this application provide a housing data processing method, device, medium, and product, which can be used in the field of data processing technology and aim to solve the above-mentioned technical problems of the prior art. By using the method of dividing data samples by community dimension and housing unit dimension, the relationship between each marginal influencing factor and housing price under different dimensions is analyzed in depth, thereby accurately determining the marginal factor coefficient of each marginal influencing factor. Using this refined coefficient determination method, the actual value of a single housing unit can be more accurately reflected, effectively improving the accuracy of the determined housing value and solving the problem of large deviation between housing value and actual transaction price caused by existing housing unit price determination methods.
[0081] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0082] Figure 1 This is a schematic diagram illustrating an application scenario of a housing data processing method provided in an embodiment of this application, such as... Figure 1 As shown, it includes: terminal 101 and server 102.
[0083] Terminal 101 is used to display the value of the target house for users to view. Server 102 is connected to terminal 101 and is used to obtain a sample dataset of houses in the area where the target house is located, and determine the value of the target house based on the sample dataset.
[0084] Figure 2 This is a flowchart illustrating a housing data processing method provided in an embodiment of this application. The executing entity in this embodiment can be...Figure 1 The server 102 in the illustrated embodiment can also be other computer-related devices, and this embodiment does not impose any particular limitation on it. For ease of description, this application embodiment uniformly describes the executing entity of a housing data processing method as a server.
[0085] like Figure 2 As shown, the method includes:
[0086] S201. Obtain the housing sample dataset of the area where the target house is located.
[0087] Specifically, the server can establish a data sharing mechanism with real estate transaction platforms, real estate agencies, and real estate management departments to collect a large amount of data related to sold or listed houses in the area where the target house is located, thus obtaining a housing sample dataset.
[0088] The housing sample dataset may include, but is not limited to, historical records of all valid listings and transactions of properties in the target property's area within a preset time period (such as the past year). Each historical record may contain basic property attribute information (such as building area, unit structure, orientation, decoration status, floor, total number of floors, building year, whether it has an elevator, etc.), geographical location information (such as the community it belongs to), price information (such as listing price per unit area, transaction price per unit area), and metadata such as timestamps.
[0089] S202. Based on the housing sample dataset, determine the sample field information that meets the preset coverage conditions.
[0090] In this embodiment of the application, the sample field information includes various factors affecting house value.
[0091] Specifically, the preset coverage condition can mean that the proportion of non-empty fields in the sample is not less than 80%. The sample field information retained after screening can constitute an effective feature set for subsequent analysis. The sample field information includes a variety of factors that have a potential impact on the value of a house, collectively referred to as house value influencing factors, such as: area, orientation, decoration level, floor location, view, noise level, rationality of the unit type, building type, and location within the community.
[0092] S203. Analyze the sample field information using a decision tree to determine the core influencing factors and multiple different marginal influencing factors.
[0093] Specifically, the server can use decision tree models, such as high-precision gradient boosting tree models like eXtreme Gradient Boosting (XGBoost), to perform feature importance analysis on sample field information and quantify the contribution of each influencing factor to the unit price of a house. Based on the analysis results, the factors with the highest influence and highest ranking (such as house area) are identified as core influencing factors. Other factors with relatively smaller but not negligible influence (such as orientation, decoration, apartment type, floor, building type, etc.) are classified into several different marginal influencing factors.
[0094] S204. Based on the housing sample dataset, determine the housing sample data corresponding to the communities where the value difference reaches the preset difference threshold as the training sample dataset.
[0095] Specifically, the server can calculate the standard deviation or coefficient of variation of the unit price of housing units in each neighborhood. If it exceeds a preset difference threshold (e.g., the standard deviation is greater than 1.5 times the regional mean), it is determined that the price differentiation within that neighborhood is significant and has modeling value, and all housing samples in that neighborhood are included in the training sample dataset. By analyzing the distribution of the standard deviation of housing prices within each neighborhood, neighborhoods with large price differences are selected.
[0096] By focusing on typical scenarios where price structures are complex and traditional mean models fail, the generalization ability of subsequent pre-built value estimation models can be improved.
[0097] S205. Perform two-dimensional bucketing on the training sample dataset to obtain data samples divided by community dimension and data samples divided by housing source dimension.
[0098] Specifically, at the neighborhood level, neighborhoods can be divided into two categories based on their overall area distribution characteristics (such as the minimum unit size or average unit size): first-area neighborhoods (i.e., large-area neighborhoods) and second-area neighborhoods (i.e., small-area neighborhoods), where the first-area neighborhood is larger than the second-area neighborhood. More specifically, the median house size in the city can be used as the boundary to divide neighborhoods into first-area and second-area neighborhoods (the median can be adjusted based on housing sample data from different cities). For example, the main distribution range of house sizes in neighborhoods within a specific city might be 40-200 square meters. The median house size in each neighborhood is 80 square meters. When the minimum unit size in a neighborhood is greater than or equal to 80 square meters, it can be identified as a first-area neighborhood; when the minimum unit size is less than 80 square meters, it can be identified as a second-area neighborhood. The definition of first-area and second-area neighborhoods can be adjusted based on different housing sample data from different cities.
[0099] In terms of housing resources, each individual housing unit within a community can be divided into five categories based on its building area according to a preset quantile: Category T, Category S, Category M, Category B, and Category L, with the housing area increasing sequentially from T, S, M, B, to L.
[0100] By using the above two-dimensional segmentation, we can obtain structured data samples for community-level segmentation and housing-level segmentation, laying the foundation for subsequent refined modeling.
[0101] S206. Perform statistical analysis on the data samples divided by community dimension and the data samples divided by housing dimension to obtain the edge factor coefficients of each edge influence factor.
[0102] Specifically, the server can calculate the conditional mean or weighted adjustment coefficient of various marginal influence factors under different buckets, thereby obtaining the marginal factor coefficient of each marginal influence factor. For example, for the orientation factor, the premium ratio of south-facing properties can be statistically analyzed in high-rise residential buildings and low-density villa communities to generate a differentiation coefficient, thus obtaining the marginal factor coefficient corresponding to the orientation factor.
[0103] S207. Based on the data sample division by community dimension, the data sample division by housing dimension, and the pre-built value estimation model, determine the model estimate of the unit price of the community sample containing the core influencing factors.
[0104] Specifically, the server can apply a pre-built value estimation model, such as a multiple linear regression model, using the net unit price after removing the influence of marginal factors as the dependent variable, and the average area, average unit price, current month's benchmark price, and area-specific labels of the community over a fixed time period (e.g., the past four months) as independent variables. The model is then fitted to the training samples, outputting a model estimate of the community's sample unit price that includes the core influencing factor (i.e., area). This model's estimate of the community's sample unit price already incorporates the community-level area factor effect, reflecting the real market's pricing elasticity based on area.
[0105] S208. Based on the model, estimate the unit price of the sample housing area and the marginal factor coefficients of each marginal influence factor to determine the target housing value.
[0106] Specifically, the server can estimate the unit price of the sample housing area based on the model, and combine this with the marginal factor coefficients of the aforementioned marginal influence factors. The estimated unit price is then comprehensively corrected through a product, thereby determining the final valuation of the target housing value and achieving high-precision, personalized housing valuation. Specifically, the formula for determining the target housing value is as follows:
[0107] Y = Model estimated unit price of the community sample × area coefficient × decoration coefficient × orientation coefficient × building type coefficient × unit type coefficient × floor coefficient.
[0108] Here, Y represents the target house value. By comprehensively considering the impact of core and peripheral influencing factors on house value, the actual value of the target house can be determined more accurately.
[0109] This embodiment provides a housing data processing method, comprising: first, acquiring a housing sample dataset of the region where the target housing is located; then, determining sample field information that meets a preset coverage condition based on the housing sample dataset, wherein the sample field information includes multiple housing value influencing factors; next, analyzing the sample field information through a decision tree to determine core influencing factors and multiple different marginal influencing factors; and, based on the housing sample dataset, determining housing sample data corresponding to communities where the value difference reaches a preset difference threshold as training sample dataset; subsequently, performing two-dimensional binning on the training sample dataset to obtain community-level segmented data samples and housing-level segmented data samples; and performing statistical analysis on the community-level segmented data samples and housing-level segmented data samples to obtain the marginal factor coefficients of each marginal influencing factor; then, based on the community-level segmented data samples, housing-level segmented data samples, and a pre-built value estimation model, determining the model-estimated community sample unit price containing the core influencing factors; finally, determining the target housing value based on the model-estimated community sample unit price and the marginal factor coefficients of each marginal influencing factor.
[0110] The following technical effects were achieved: Based on feature importance analysis, the area factor was identified as the core variable with the most significant impact on the unit price of houses. A two-dimensional segmentation mechanism was then introduced: at the community level, communities were divided into first-area communities and second-area communities based on their overall area distribution; at the housing unit level, individual housing units were further subdivided into five categories based on area. Finally, for different community types, data samples were segmented based on the community and housing unit dimensions, and a pre-built value estimation model was established. The estimated unit price of community samples was determined by removing interference from various marginal factors. Combined with the marginal factor coefficients of each marginal factor, the target house value was determined. This approach breaks through the limitations of traditional reliance on static expert rules, achieving an evolution from determining marginal factor coefficients using a unified standard to dynamic modeling at the community level, significantly improving the consistency between house valuation results and actual transaction prices. By comprehensively considering the impact of multiple influencing factors on house value, the accuracy of house valuation was effectively improved, providing strong support for real estate transactions and market supervision.
[0111] In one possible implementation, the sample field information is analyzed using a decision tree to determine the core influencing factor and multiple different marginal influencing factors. This includes: using a preset analysis tool to analyze the sample field information using the decision tree to obtain multiple housing value influencing factors and their corresponding influence scores; performing numerical comparison processing based on each housing value influencing factor and its corresponding influence score to determine the housing value influencing factor with the largest influence score as the core influencing factor, and identifying the other housing value influencing factors besides the core influencing factor as marginal influencing factors.
[0112] Specifically, the server first uses sample field information as input features, imports it into a pre-defined decision tree analysis model, such as XGBoost, Lightweight Gradient Boosting Machine (LightGBM), or Random Forest, and uses its built-in feature importance assessment mechanism (such as quantification methods based on information gain or number of splits) to score each housing value influencing factor, obtaining a feature importance score for each factor. This score characterizes the degree to which the corresponding housing value influencing factor contributes to the housing unit price prediction result.
[0113] The server can then perform numerical comparisons of all housing value influencing factors and their corresponding influence scores. Specifically, the housing value influencing factor with the highest influence score can be selected as the core influencing factor. The core influencing factor plays a dominant role in housing value and usually has the strongest explanatory power for price changes (for example, in most cities, building area is usually identified as a core influencing factor).
[0114] Finally, the server can categorize all other factors influencing housing value besides the core influencing factors into marginal influencing factors. Although the individual influence of these marginal influencing factors is not as great as that of the core influencing factors, their combined effect still has a significant corrective effect on the final housing price.
[0115] In one possible implementation, the training sample dataset is subjected to two-dimensional binning to obtain data samples divided by the community dimension and data samples divided by the housing source dimension. This includes: dividing the training sample dataset by the median area of the community dimension to obtain a first area community dataset and a second area community dataset as community dimension data samples, wherein the first area is larger than the second area; and dividing the training sample dataset by the housing source dimension by the housing area to obtain multiple sets of housing sample data of different housing types as housing source dimension data samples.
[0116] Specifically, firstly, at the community level, the median area (or representative statistics such as the smallest unit size or average area) of all housing units within each community is calculated, and all communities are then divided based on this indicator. Specifically, all communities are compared with a preset city-level area threshold (e.g., the global median of the area of all communities in that city).
[0117] If the median area of a specific community is greater than or equal to the preset city-level area threshold, it will be classified into the first area community dataset (i.e., large area community).
[0118] If the median area of a specific community is less than the preset city-level area threshold, then the community is classified into the second area community dataset (i.e., small area community).
[0119] The resulting datasets of the first and second area neighborhoods together constitute a neighborhood-dimension segmentation data sample, used to reflect the structural differences in the relationship between price and area for neighborhoods of different sizes.
[0120] Secondly, at the housing unit level, for each individual unit within a community, multi-level binning can be performed based on its actual building area. Specifically, based on the distribution quantiles of housing unit area within the community (e.g., 20%, 40%, 60%, 80% quantiles), housing units are divided into five categories: T-type, S-type, M-type, B-type, and L-type, with housing unit area increasing sequentially from T to L. This division yields multiple sets of housing sample data of different housing unit types, serving as data samples for housing unit dimension segmentation.
[0121] By using dual-dimensional binning, we can achieve refined stratification from macro (overall community positioning) to micro (individual property attributes), providing a structured data foundation for determining the edge factor coefficients of each edge influencing factor for different scenarios.
[0122] In one possible implementation, statistical analysis is performed on the data samples divided by the community dimension and the data samples divided by the housing source dimension to obtain the edge factor coefficients of each edge influencing factor. This includes: performing a correlation analysis of housing area and unit price based on the data samples divided by the community dimension to obtain the correlation analysis results of housing area and price under the community dimension; and performing statistical calculations based on the correlation analysis results of housing area and price under the community dimension and the data samples divided by the housing source dimension to obtain the area coefficient, decoration coefficient, orientation coefficient, building type coefficient, unit type coefficient, and floor coefficient as the edge factor coefficients of each edge influencing factor.
[0123] Specifically, firstly, the server can divide the data samples based on the neighborhood dimension (i.e., the first area neighborhood dataset and the second area neighborhood dataset) and conduct correlation analysis between housing area and unit price separately. Specifically, within each neighborhood category, the average transaction unit price corresponding to different area bins (such as T-type housing, S-type housing, M-type housing, B-type housing, and L-type housing) can be calculated, and the scatter distribution trend of area and unit price can be fitted to obtain the correlation analysis results between area and price at the neighborhood dimension.
[0124] The correlation analysis of area and price under this community dimension reveals that in the first area community, the impact of area on unit price tends to be moderate, and the correlation between house area and price is relatively small.
[0125] In the second-largest residential areas, smaller units typically exhibit a significant price premium per unit, with the area and price of the house tending to be negatively correlated; the smaller the area, the higher the price per unit.
[0126] It should be noted that the housing area, as a core influencing factor, has already been embedded in the model estimation of the unit price of the sample community through the pre-built value estimation model (such as the linear regression model). Therefore, it will not be treated as a marginal influencing factor in this step.
[0127] Subsequently, the server can combine the above correlation analysis results and data samples segmented by housing dimensions to perform conditional statistical calculations on other housing attribute fields besides area, in order to generate the marginal factor coefficients of each marginal influencing factor. Specifically, for decoration status, within the same community and similar area bins, the average unit price ratio of fully furnished apartments to partially furnished / unfurnished apartments is calculated as the decoration coefficient.
[0128] Regarding orientation, the average premium of properties with mainstream orientations such as south and east relative to those with north or west orientations is calculated to generate an orientation coefficient.
[0129] Similarly, for factors such as building type (e.g., slab building, tower building, villa), unit structure (e.g., whether it is square, whether there is a separation between active and quiet areas), and floor (e.g., low zone, middle zone, high zone, top floor), under the premise of controlling the main variables such as area and community, the conditional mean ratio or weighted adjustment coefficient is calculated respectively to obtain the corresponding building type coefficient, unit type coefficient and floor coefficient.
[0130] Finally, the above coefficients are used as the marginal factor coefficients of each marginal influence factor, which are then used to refine the estimated unit price of the community sample in the model, so as to achieve a high-precision valuation of the house value.
[0131] In one possible implementation, based on data samples segmented by community dimension, data samples segmented by housing dimension, and a pre-built value estimation model, the model-estimated unit price of the community sample is determined, including: determining the average area and average unit price within a first fixed time period based on data samples segmented by community dimension and data samples segmented by housing dimension; determining the benchmark price within a second fixed time period based on data samples segmented by community dimension and data samples segmented by housing dimension; determining binned composite data based on data samples segmented by community dimension and data samples segmented by housing dimension; determining the sample area based on data samples segmented by community dimension and data samples segmented by housing dimension; and determining the model-estimated unit price of the community sample, including the core influencing factors, based on the average area and average unit price within the first fixed time period, the benchmark price within the second fixed time period, the binned composite data, the sample area, and the pre-built value estimation model.
[0132] Specifically, firstly, the server can divide data samples based on the neighborhood dimension (including the first and second area neighborhoods) and the corresponding housing unit dimension (i.e., various housing units after being binned by area), and extract key feature variables. This can include:
[0133] Determine the average area and average price per unit within a first fixed time period. For example, you can calculate the average building area and average price per unit of all valid transactions or listings in each neighborhood over the past four months to characterize the recent market supply and demand and product structure characteristics of the neighborhood. The average area within the first fixed time period reflects the size of the neighborhood's main unit types recently, while the average price per unit represents the overall price level of the neighborhood recently. The first fixed time period can be the past four months, the past six months, or the past three months, etc., without specific restrictions.
[0134] Determine the benchmark price for the second fixed time period. For example, you can obtain the benchmark price of the property in the current month published on the platform for the same building or area, as an external market anchor to reflect the macro pricing trend of the region. The second fixed time period can be the past month, the past two months, or the past half month, etc., and there are no specific restrictions here.
[0135] Identify the binned composite data. For example, the above two-dimensional binning results can be encoded into structured features, including neighborhood type labels (e.g., 0 represents the first area neighborhood, 1 represents the second area neighborhood) and the area bin category to which the property belongs (e.g., T-class property = 0, S-class property = 1, M-class property = 2, B-class property = 3, L-class property = 4), thus forming a combined categorical variable to capture nonlinear interaction effects and encode the combined features of neighborhood area type and property area bin category.
[0136] Determine the sample area; for example, obtain the actual building area of the property to be valued, as the original input for the core influencing factor.
[0137] Then, the server can input the average area and average unit price within the first fixed time period, the benchmark price within the second fixed time period, the composite data of the buckets, and the sample area into the pre-built value estimation model to determine the model estimate of the sample unit price of the community containing the core influencing factors.
[0138] The pre-built value estimation model can be a multiple linear regression model, with the following specific form:
[0139]
[0140] Where y is the model-estimated unit price of the sample community including core influencing factors, which refers to the net unit price after removing the interference of marginal influencing factors. x1 is the average area in the first fixed time period, x2 is the average unit price in the first fixed time period, x3 is the benchmark price in the second fixed time period, x4 is the binned composite data, and x5 is the sample area. to All model parameters, including ε, can be estimated using the least squares method. The specific estimation method is based on existing technology and will not be elaborated here.
[0141] After fitting the model parameters using the least squares method, the pre-built value estimation model can be used to predict the target sample. The output is the model-estimated unit price of the cell sample, which includes the core influencing factors. The model-estimated unit price of the cell sample has embedded the cell-level area elasticity relationship, which can serve as the basis for subsequent fusion and edge factor correction.
[0142] In one possible implementation, before determining the sample field information that meets the preset coverage conditions based on the house sample dataset, the method further includes: performing noise reduction processing on the house sample dataset to obtain an interference-resistant house sample dataset.
[0143] Specifically, firstly, the server can perform data quality checks on each record in the original housing sample dataset, removing samples with obvious anomalies or invalid values. For example, it can remove properties with a building area of less than 10 square meters or greater than 1000 square meters (beyond the reasonable residential range). It can also remove extreme price samples with a unit price lower than the regional minimum guaranteed price or higher than the reasonable market upper limit (e.g., exceeding three standard deviations of the average price in the same community). Finally, it can remove records with missing or incorrectly formatted key fields (such as area, price, and community name).
[0144] Secondly, redundant data such as duplicate listings or multiple transactions of the same property can be deduplicated, retaining the latest or most representative records to avoid sample bias.
[0145] Furthermore, for abnormal fluctuations in the time dimension, a sliding window mechanism can be used to identify and filter noise behaviors such as short-term order brushing and false pricing.
[0146] After the above cleaning and filtering, the obtained anti-interference housing sample dataset can significantly improve the data integrity, rationality and representativeness, providing a high-quality data input foundation for subsequent field coverage screening and feature importance analysis.
[0147] Figure 3 This is a schematic diagram of a housing data processing device provided in an embodiment of this application. Figure 3 As shown, the housing data processing device includes: an acquisition module 310, an influence factor determination module 320, a training data determination module 330, a coefficient determination module 340, a core unit price determination module 350, and a housing value determination module 360.
[0148] The acquisition module 310 is used to acquire a sample dataset of houses in the area where the target house is located;
[0149] The influencing factor determination module 320 is used to determine the sample field information that meets the preset coverage conditions based on the housing sample dataset, wherein the sample field information includes a variety of housing value influencing factors;
[0150] The impact factor determination module 320 is also used to analyze sample field information through decision trees to determine core impact factors and multiple different marginal impact factors;
[0151] The training data determination module 330 is used to determine the housing sample data corresponding to the community whose value difference reaches a preset difference threshold as the training sample dataset based on the housing sample dataset.
[0152] The training data determination module 330 is also used to perform two-dimensional bucketing on the training sample dataset to obtain community-dimension segmented data samples and housing-dimension segmented data samples.
[0153] The coefficient determination module 340 is used to perform statistical analysis on the data samples divided by the community dimension and the data samples divided by the housing dimension to obtain the marginal factor coefficients of each marginal influence factor.
[0154] The core unit price determination module 350 is used to divide data samples according to the community dimension, the housing dimension, and the pre-built value estimation model to determine the model estimate of the community sample unit price containing core influencing factors.
[0155] The housing value determination module 360 is used to determine the target housing value based on the estimated unit price of the sample housing area and the marginal factor coefficients of each marginal influencing factor according to the model.
[0156] In one possible implementation, the influence factor determination module 320 is also used to obtain multiple house value influence factors and the influence scores corresponding to each house value influence factor by using a preset analysis tool on the sample field information through a decision tree.
[0157] The impact factor determination module 320 is also used to perform numerical comparison processing based on each housing value impact factor and the impact score corresponding to each housing value impact factor, to obtain the housing value impact factor corresponding to the largest impact score as the core impact factor, and to determine the other housing value impact factors besides the core impact factor as marginal impact factors.
[0158] In one possible implementation, the training data determination module 330 is further configured to perform area median division processing on the training sample dataset in the cell dimension to obtain a first area cell dataset and a second area cell dataset as cell dimension division data samples, wherein the first area is larger than the second area.
[0159] The training data determination module 330 is also used to perform house area binning on the training sample dataset in the house source dimension to obtain multiple sets of house sample data with different house source types as house source dimension partitioning data samples.
[0160] In one possible implementation, the coefficient determination module 340 is also used to divide the data samples according to the community dimension to perform a correlation analysis of housing area and unit price, and obtain the correlation analysis results of housing area and price under the community dimension.
[0161] The coefficient determination module 340 is also used to perform statistical calculations based on the correlation analysis results between housing area and price under the community dimension and the data samples divided by the housing source dimension, to obtain the area coefficient, decoration coefficient, orientation coefficient, building type coefficient, unit type coefficient and floor coefficient as the edge factor coefficients of each edge influencing factor.
[0162] In one possible implementation, the core unit price determination module 350 is also used to divide the data samples according to the community dimension and the housing dimension to determine the average area and average unit price within a first fixed time period.
[0163] The core unit price determination module 350 is also used to divide data samples according to the community dimension and the housing dimension to determine the benchmark price in the second fixed time period;
[0164] The core unit price determination module 350 is also used to divide data samples according to the community dimension and the housing dimension to determine the binned composite data;
[0165] The core unit price determination module 350 is also used to divide data samples according to the community dimension and the housing dimension to determine the sample area;
[0166] The core unit price determination module 350 is also used to determine the model estimated unit price of the community sample based on the average area and average unit price in the first fixed time period, the benchmark price in the second fixed time period, the binned composite data, the sample area and the pre-built value estimation model, which includes core influencing factors.
[0167] In one possible implementation, in the housing value determination module 360, the calculation formula for determining the target housing value is as follows, based on the estimated unit price of the community sample and the marginal factor coefficients of each marginal influence factor:
[0168] Y = Model estimated unit price of the community sample × area coefficient × decoration coefficient × orientation coefficient × building type coefficient × unit type coefficient × floor coefficient.
[0169] In one possible implementation, the device further includes a preprocessing module;
[0170] The preprocessing module is used to denoise the house sample dataset to obtain an interference-resistant house sample dataset.
[0171] The housing data processing device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0172] Figure 4 This is a schematic diagram of the structure of an electronic device provided in this application. Figure 4 As shown, the electronic device provided in this embodiment includes at least one processor 410 and a memory 420. The electronic device also includes a communication component 430. The processor 410, memory 420, and communication component 430 are connected via a bus 440.
[0173] In the specific implementation process, at least one processor 410 executes computer execution instructions stored in memory 420, causing at least one processor 410 to execute a house data processing method as executed on the electronic device side as described above.
[0174] The specific implementation process of processor 410 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0175] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0176] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0177] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0178] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0179] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0180] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0181] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0182] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0183] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0184] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0185] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0186] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0187] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for processing housing data, characterized in that, include: Obtain a sample dataset of houses in the area where the target house is located; Based on the housing sample dataset, determine the sample field information that meets the preset coverage conditions, wherein the sample field information includes multiple housing value influencing factors; The sample field information is analyzed using a decision tree to determine the core influencing factors and multiple different marginal influencing factors; Based on the housing sample dataset, housing sample data corresponding to communities where the value difference reaches a preset difference threshold are determined as training sample datasets. The training sample dataset is subjected to two-dimensional bucketing to obtain data samples divided by community dimension and data samples divided by housing source dimension. Statistical analysis was performed on the data samples divided by the community dimension and the data samples divided by the housing dimension to obtain the edge factor coefficients of each edge influence factor; Based on the data samples divided by the community dimension, the data samples divided by the housing dimension, and the pre-built value estimation model, the model estimate of the unit price of the community sample, which includes the core influencing factors, is determined. The target house value is determined by estimating the sample unit price of the community and the edge factor coefficients of each edge influence factor based on the model.
2. The method according to claim 1, characterized in that, The analysis of the sample field information using a decision tree determines the core influencing factor and multiple distinct marginal influencing factors, including: By using a decision tree and a preset analysis tool on the sample field information, various factors affecting house value and the corresponding impact scores of each factor are obtained. The housing value influencing factors and their corresponding influence scores are compared numerically to identify the housing value influencing factor with the largest influence score as the core influencing factor, and the other housing value influencing factors besides the core influencing factor are identified as marginal influencing factors.
3. The method according to claim 1, characterized in that, The step of performing two-dimensional bucketing on the training sample dataset to obtain data samples divided by community dimension and data samples divided by housing unit dimension includes: The training sample dataset is divided into cell-level data samples by the median area, resulting in a first area cell dataset and a second area cell dataset, where the first area is larger than the second area. The training sample dataset is binned by house area in the house source dimension to obtain multiple sets of house sample data with different house types as house source dimension partitioning data samples.
4. The method according to claim 1, characterized in that, The statistical analysis of the data samples segmented by the community dimension and the data samples segmented by the housing resource dimension yields the marginal factor coefficients of each marginal influence factor, including: Based on the data samples divided according to the community dimension, a correlation analysis of housing area and unit price was performed to obtain the results of the correlation analysis between housing area and price under the community dimension. Based on the correlation analysis results of housing area and price under the community dimension and the data sample divided by the housing source dimension, statistical calculations are performed to obtain the area coefficient, decoration coefficient, orientation coefficient, building type coefficient, unit type coefficient and floor coefficient as the edge factor coefficients of each edge influencing factor.
5. The method according to claim 1, characterized in that, The step of determining the estimated unit price of the community sample based on the data sample segmentation according to the community dimension, the data sample segmentation according to the housing resource dimension, and the pre-built value estimation model, including the model estimation of core influencing factors, includes: Based on the data samples divided by the community dimension and the data samples divided by the housing dimension, the average area and average unit price within the first fixed time period are determined. Based on the data samples divided by the community dimension and the data samples divided by the housing dimension, the benchmark price for the second fixed time period is determined. Based on the data samples divided according to the community dimension and the data samples divided according to the housing dimension, the binned composite data is determined; The sample area is determined by dividing the data samples according to the community dimension and the housing dimension. Based on the average area and average unit price within the first fixed time period, the benchmark price within the second fixed time period, the binned composite data, the sample area, and the pre-built value estimation model, the model-estimated sample unit price of the community, which includes core influencing factors, is determined.
6. The method according to claim 1, characterized in that, The formula for determining the target house value based on the estimated unit price of the community sample and the edge factor coefficients of each edge influence factor according to the model is as follows: Y = Model estimated unit price of the community sample × area coefficient × decoration coefficient × orientation coefficient × building type coefficient × unit type coefficient × floor coefficient.
7. The method according to claim 1, characterized in that, Before determining the sample field information that meets the preset coverage conditions based on the housing sample dataset, the process also includes: The house sample dataset is denoised to obtain an interference-resistant house sample dataset.
8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.