House price automatic evaluation data anomaly detection method and system
By dividing urban geographic space into geographic units, calculating property holding costs and market price change indicators, quantifying the degree of contradiction, identifying and expanding abnormal areas, and adjusting the valuation process, this method solves the problem of difficulty in identifying complex data anomalies in existing technologies, and improves the accuracy and reliability of housing price valuation.
Patent Information
- Application Number
- CN202511099943.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing automated housing price valuation systems struggle to identify complex data anomaly patterns that are semantically contradictory but logically coherent and triggered by the same physical event, leading to systemic and clustered biases in the valuation system.
The city's geographic space is divided into multiple geographic units. Data is collected based on the geographic location of the properties. Holding costs and market price change indicators are calculated to determine whether they show an upward trend and reach a preset threshold within the same analysis period. The degree of contradiction is quantified, overall abnormal areas are identified and expanded, and the valuation process is adjusted based on the abnormality markers.
Effectively identify and handle complex data anomalies to improve the accuracy and reliability of housing price valuation and avoid systematic biases.
Smart Images

Figure CN120996840A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of data anomaly detection, specifically to a method and system for detecting data anomalies in automatic house price valuation. Background Technology
[0002] In modern urban management and financial services, automated property valuation systems have become indispensable tools. These systems collect and integrate massive amounts of property information, providing crucial value for credit approval and real estate transactions. Their data systems typically consist of two main parts: static information describing the characteristics of the property itself and transaction information reflecting market dynamics. This raw information is captured by the system from various data sources, processed initially, and stored in a core valuation database to drive subsequent valuation calculations.
[0003] In practical applications, when cities face external physical events such as large-scale infrastructure construction, even if these events have a clear spatial impact range, they may trigger signal flows with semantically contradictory data characteristics but logically common origins within the same geographical area. Existing anomaly detection methods often struggle to identify such complex anomaly patterns. Summary of the Invention
[0004] The purpose of this invention is to address the aforementioned shortcomings by proposing a method and system for detecting anomalies in automatic housing price valuation data.
[0005] The present invention adopts the following technical solution:
[0006] An automatic housing price valuation data anomaly detection method includes the following steps: dividing the urban geographic space into multiple geographic units, and aggregating housing data into the corresponding geographic units based on the geographical location of the properties; for each geographic unit, calculating the holding cost change index and market price change index of the properties within that geographic unit; determining whether the holding cost change index and market price change index of the properties within the geographic unit simultaneously show an upward trend within the same analysis period, and whether their respective changes reach preset thresholds; based on the judgment results, quantifying the degree of contradiction when the holding cost change index and market price change index of the properties within the geographic unit simultaneously show an upward trend within the same analysis period; based on the degree of contradiction, identifying and expanding the overall abnormal area affected by the contradiction; adding anomaly tags to the housing data within the overall abnormal area, and adjusting the housing valuation process according to the anomaly tags.
[0007] The above scheme can effectively identify and process complex data anomaly patterns that are semantically contradictory but logically coherent and caused by the same physical event, thereby avoiding systematic and clustered biases in the valuation system and improving the accuracy and reliability of housing price valuation.
[0008] To further address this issue, this application also proposes a step for determining whether the holding cost change index and market price change index of real estate within a geographic unit simultaneously show an upward trend within the same analysis period, and whether their respective changes reach a preset threshold. This step includes: continuously recording the market price change index of real estate within the geographic unit; when the market price change index of real estate within the geographic unit shows an upward trend and reaches a preset threshold, generating a price increase record for real estate within the geographic unit, which includes the time of the price increase; when the holding cost change index of real estate within the geographic unit shows an upward trend and reaches a preset threshold, retrieving a price increase record for real estate within the geographic unit within a preset time window; if a price increase record for real estate within the geographic unit exists within the preset time window, then it is determined that the holding cost change index and market price change index of real estate within the geographic unit simultaneously show an upward trend within the same analysis period.
[0009] To improve the solution, this application also proposes that the steps for calculating the change index of holding costs of real estate within a geographic unit include: identifying the management entity to which the real estate within the geographic unit belongs; aggregating the holding cost data of the management entity to obtain unified holding cost data at the management entity level; and using the unified holding cost data at the management entity level when calculating the change index of holding costs of real estate within the geographic unit.
[0010] To improve the solution, this application also proposes a step for quantifying the degree of contradiction when the indicators of changes in holding costs and market prices of real estate within a geographic unit simultaneously show an upward trend within the same analysis period, based on the judgment results. This step includes: obtaining market activity information for the geographic unit; obtaining information on the types of real estate within the geographic unit; obtaining market trend information for the macroeconomic cycle; selecting contradiction quantification parameters suitable for the geographic unit from a preset parameter set based on the market activity information, the types of real estate within the geographic unit, and the market trend information for the macroeconomic cycle; and using the contradiction quantification parameters, calculating the degree of contradiction when the indicators of changes in holding costs and market prices of real estate within the geographic unit simultaneously show an upward trend within the same analysis period, based on the judgment results.
[0011] To improve the solution, this application also proposes a step for identifying and expanding the overall anomalous region affected by contradictions based on the degree of contradiction: identifying geographical units whose contradiction degree reaches a preset contradiction threshold based on the degree of contradiction, and marking them as anomalous candidate units; obtaining the geographical attribute information of the anomalous candidate units; obtaining the contradiction pattern characteristics of the anomalous candidate units; merging non-adjacent anomalous candidate units on the geographical unit according to the geographical attribute information and contradiction pattern characteristics of the anomalous candidate units to form discrete anomalous regions; expanding adjacent anomalous candidate units on the geographical unit according to the geographical adjacency relationship of the anomalous candidate units to form continuous anomalous regions; and merging discrete anomalous regions and continuous anomalous regions to form an overall anomalous region.
[0012] To improve the solution, this application also proposes a step for merging non-adjacent anomalous candidate units on a geographic unit to form a discrete anomalous region based on the geographic attribute information and contradictory pattern characteristics of the anomalous candidate units. This step includes: analyzing the geographic attribute information of the anomalous candidate units to identify units sharing geographic features; analyzing the contradictory pattern characteristics of the anomalous candidate units to identify units with similar contradictory patterns; and merging non-adjacent anomalous candidate units on a geographic unit based on units with shared geographic features and units with similar contradictory patterns to form a discrete anomalous region.
[0013] To improve the solution, this application also proposes that the steps for adjusting the property valuation process based on anomaly markers include: selecting a valuation sub-model corresponding to the anomaly marker; and using the valuation sub-model to perform valuation processing on the property data within the overall anomaly area, thereby adjusting the property valuation process.
[0014] To improve the solution, this application also proposes a step for selecting an evaluation sub-model corresponding to an anomaly marker based on the anomaly marker, including: parsing the anomaly marker to obtain the anomaly features contained in the anomaly marker; identifying multiple evaluation sub-models that can handle the anomaly features based on the anomaly features; ranking the multiple evaluation sub-models according to the degree of adaptation of each evaluation sub-model to the anomaly features; and selecting the evaluation sub-model with the highest priority in the ranking results as the evaluation sub-model corresponding to the anomaly marker.
[0015] To improve the solution, this application also proposes a step for ranking multiple valuation sub-models based on the degree of adaptation of each valuation sub-model to the abnormal features, including: acquiring the abnormal features; establishing a historical database to record the historical performance data of each valuation sub-model when handling different abnormal features; retrieving historical performance data of valuation sub-models related to the abnormal features from the historical database based on the abnormal features; evaluating the degree of adaptation of each valuation sub-model to the abnormal features based on the retrieval results; and ranking the multiple valuation sub-models based on the evaluation results.
[0016] To improve the solution, this application also proposes an automatic housing price valuation data anomaly detection system, applied to an automatic housing price valuation data anomaly detection method. The system includes: a geographic unit processing module, which divides the urban geographic space into multiple geographic units and aggregates property data to the corresponding geographic units based on the property's geographical location; an indicator calculation module, which calculates the holding cost change indicator and market price change indicator for each geographic unit; a judgment module, used to determine whether the holding cost change indicator and market price change indicator for properties within a geographic unit simultaneously show an upward trend within the same analysis period, and whether their respective changes reach preset thresholds; a quantification module, based on the judgment results, quantifying the degree of contradiction when the holding cost change indicator and market price change indicator for properties within a geographic unit simultaneously show an upward trend within the same analysis period; an anomaly area processing module, which identifies and expands the overall anomaly area affected by the contradiction based on the degree of contradiction; and an adjustment module, used to attach anomaly markers to the property data within the overall anomaly area and adjust the property valuation process based on the anomaly markers.
[0017] To further understand the features and technical content of the present invention, please refer to the following detailed description and drawings of the present invention. However, the drawings provided are for reference and illustration only and are not intended to limit the present invention. Attached Figure Description
[0018] Figure 1 This is a flowchart of a method for detecting anomalies in automatic house price valuation data according to the present invention;
[0019] Figure 2 This is a schematic diagram of the structure of an automatic housing price valuation data anomaly detection system according to the present invention. Detailed Implementation
[0020] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention. Furthermore, the accompanying drawings of the present invention are for simple illustrative purposes only and are not depictions of actual dimensions; this is stated in advance. The following embodiments will further describe the relevant technical content of the present invention in detail, but the disclosed content is not intended to limit the scope of protection of the present invention.
[0021] This embodiment provides a method and system for detecting anomalies in automatic house price valuation data, combined with... Figure 1 and Figure 2 As shown.
[0022] refer to Figure 1An automatic housing price valuation data anomaly detection method is proposed, comprising the following steps: dividing the urban geographic space into multiple geographic units, and aggregating housing data into the corresponding geographic units according to the geographical location of the properties; for each geographic unit, calculating the holding cost change index and market price change index of the properties within that geographic unit; determining whether the holding cost change index and market price change index of the properties within the geographic unit simultaneously show an upward trend within the same analysis period, and whether their respective changes reach preset thresholds; based on the judgment results, quantifying the degree of contradiction when the holding cost change index and market price change index of the properties within the geographic unit simultaneously show an upward trend within the same analysis period; based on the degree of contradiction, identifying and expanding the overall abnormal area affected by the contradiction; attaching anomaly tags to the housing data within the overall abnormal area, and adjusting the housing valuation process according to the anomaly tags.
[0023] Geographic units refer to the logical or physical division of urban space. This can be achieved through grid division, administrative divisions, community boundaries, or custom areas. For example, a city map can be overlaid with square or hexagonal grids, or division can be based on natural boundaries such as streets and rivers. The primary purpose is to achieve spatial aggregation and localized management of massive amounts of real estate data, laying the foundation for subsequent regional analysis. Holding cost change indicators are quantitative values reflecting the changes in various expenses during property ownership over time. These can be calculated using the rate of change or absolute difference of data related to property management fees, property taxes, maintenance costs, and loan interest. For example, comparing the average increase in property management fees at different points in time helps capture changes in the actual economic burden borne by property owners. Market price change indicators are quantitative values reflecting the changes in the transaction or listing prices of real estate in the market over time. These can be calculated using the rate of change or absolute difference of data such as historical transaction prices, current listing prices, market valuations, or regional average housing prices. For example, monitoring the average increase in listing prices within a specific area helps reflect the dynamic adjustment of market expectations for property value. Simultaneous upward trends within the same analysis period, with each indicator reaching a preset threshold, refer to a situation where, within a defined time span, both the property holding cost change indicator and the market price change indicator exhibit a sustained upward trend, and their respective growth rates reach or exceed preset numerical limits. This can be determined using methods such as time series analysis, trend fitting, or threshold comparison. For example, a monthly increase exceeding 5% might be defined as an upward trend reaching a threshold. This primarily aims to identify specific market signals of abnormally simultaneous increases in costs and prices, serving as a preliminary basis for identifying potential contradictions. The degree of contradiction refers to the extent to which the property holding cost change indicator and the market price change indicator within a quantitative geographical unit exhibit upward trends simultaneously within the same analysis period, indicating a logical inconsistency or deviation from economic laws. This can be quantified using methods such as weighted scoring, deviation calculation, or anomaly scores based on machine learning models. For example, it can be calculated based on the ratio of cost to price increases or the degree of deviation from historical normal patterns. This primarily aims to distinguish different degrees of anomalies, providing a more refined basis for subsequent identification of abnormal areas. Overall anomalous areas refer to a set of spatially clustered anomalous areas formed by identifying and expanding geographical units affected by contradictions. These anomalous areas can be identified and expanded using clustering algorithms based on geographical proximity, attribute similarity, or contradiction pattern characteristics. For example, adjacent geographical units with a high degree of contradiction can be merged, or units that are not adjacent but have similar anomalous patterns can be grouped together. The main purpose is to centrally process real estate data affected by the same external event and avoid bias caused by isolated analysis.Anomaly markers are identifiers attached to real estate data within an overall anomaly area to indicate that the data is abnormal and requires special handling. They can be represented by Boolean flags, anomaly type codes, or anomaly severity levels, such as being marked as "cost-price contradiction anomaly" or "subway construction impact area". Their main purpose is to classify and manage abnormal data and guide subsequent real estate valuation processes to make adaptive adjustments.
[0024] This application's solution systematically identifies and handles anomalies in automated property valuation data through a series of logically progressive steps. First, to effectively manage and locally analyze massive amounts of property data, the urban geographic space is divided into multiple geographic units, and property data is aggregated into corresponding geographic units based on the property's geographical location. This spatial aggregation allows subsequent indicator calculations and contradiction analysis to focus on property sets with similar geographic attributes, thereby improving the relevance of the analysis. Based on this, for each geographic unit, the system calculates the corresponding indicators for changes in holding costs and market prices. The indicator for changes in holding costs reflects the dynamics of property holding burdens, while the indicator for changes in market expectations of property value. The parallel calculation of these two indicators provides a data foundation for identifying potential contradictions. Furthermore, the system determines whether the indicators for changes in holding costs and market prices within a geographic unit simultaneously show an upward trend within the same analysis period, and whether their respective changes reach preset thresholds. This judgment step is crucial for identifying abnormal patterns; it aims to capture phenomena that should not occur simultaneously under normal economic logic, namely, a simultaneous and significant increase in costs and prices. By setting preset thresholds, normal market fluctuations can be effectively filtered out, ensuring that identified trends have actual anomaly significance. Based on this judgment, the system quantifies the degree of contradiction when the indicators of changes in property holding costs and market price changes within a geographical unit simultaneously show an upward trend within the same analysis period. This quantification process allows for differentiation of the severity of anomalies, providing a refined basis for subsequent decision-making. The calculation of the degree of contradiction considers multiple factors, ensuring the accuracy of the quantification results. Subsequently, based on the quantified degree of contradiction, the system identifies and expands the overall anomaly area affected by the contradiction. This step not only identifies individual anomalous geographical units, but more importantly, by considering spatial proximity and the similarity of contradiction patterns, it merges and expands discrete or continuous anomalous units affected by the same external event, forming a complete, spatially clustered anomaly area. This regional identification avoids misjudging isolated data points and ensures comprehensive coverage of systemic biases. Finally, anomaly tags are added to the property data within the overall anomaly area, and the property valuation process is adjusted based on these tags. By attaching anomaly markers, the system can perform special processing on data identified as anomalous, such as selecting specific valuation models or adjusting valuation parameters, thereby correcting valuation biases caused by data inconsistencies. This adaptive valuation process adjustment ensures the accuracy and robustness of house price valuations in complex market environments.
[0025] In some preferred embodiments, firstly, when dividing the urban geographic space into multiple geographic units, a grid-based method using Geographic Information System (GIS) can be employed. For example, the city map can be divided into square grids with sides of 500 meters, each grid representing a geographic unit. Property data is then automatically aggregated into its respective grid unit based on its latitude and longitude coordinates. For each geographic unit, the holding cost change indicator can be specifically calculated by monitoring the monthly growth rate of the average property management fee and property tax within that unit. For example, if the average property management fee within a geographic unit increases by more than 3% month-on-month for three consecutive months, its holding cost is considered to be on an upward trend. The market price change indicator can be specifically calculated by monitoring the monthly growth rate of the average listing price or historical transaction price of properties within that unit. For example, if the average listing price increases by more than 5% month-on-month for three consecutive months, its market price is considered to be on an upward trend. The preset threshold can be set at a monthly holding cost growth rate of 2% and a monthly market price growth rate of 4%. When the system determines that both the holding cost change indicator and the market price change indicator of properties within a geographic unit show an upward trend within the same analysis period (e.g., three consecutive months), and their respective changes reach preset thresholds, it triggers a contradiction quantification. The quantification of contradiction level can be achieved through a weighted function that comprehensively considers factors such as the difference between the holding cost growth rate and the market price growth rate, the market activity of the geographic unit (e.g., transaction volume or listing volume), property type (e.g., residential or commercial), and macroeconomic cycle (e.g., whether it is in an economic upswing or downswing). The function outputs a contradiction score between 0 and 100, with higher scores indicating greater contradiction. Based on the contradiction level, the system identifies geographic units whose contradiction scores reach a preset contradiction threshold (e.g., 60 points) and marks them as anomalous candidate units. Subsequently, the system analyzes the geographic attribute information (e.g., whether it is close to large infrastructure projects) and contradiction pattern characteristics of these anomalous candidate units (e.g., whether the cost increase is due to property fee adjustments or tax increases, and whether the price increase is due to the school district effect or transportation benefits), and then merges and expands accordingly. For example, if multiple non-adjacent anomaly candidate units are all located within the planning area of the same newly built subway line and have similar conflict patterns, they are grouped into discrete anomaly regions. Simultaneously, if adjacent anomaly candidate units all have high conflict scores, they are expanded into continuous anomaly regions. Finally, these discrete and continuous regions are merged to form an overall anomaly region. Finally, anomaly markers are added to the property data within the overall anomaly region, such as "Subway Construction Impact Area – Cost-Price Conflict." Based on this anomaly marker, the system automatically adjusts the property valuation process. Specifically, the system selects the valuation sub-model corresponding to the anomaly feature based on the anomaly marker.For example, for properties in the "subway construction impact area - cost-price contradiction", the system may prioritize the valuation sub-model based on the cost approach or income approach, and conduct more rigorous screening or adjustment of comparable cases in the market comparison approach, in order to avoid valuation deviations caused by overly high market expectations or underestimation of holding costs.
[0026] This application further proposes a step for determining whether the holding cost change index and market price change index of real estate within a geographic unit simultaneously show an upward trend within the same analysis period, and whether their respective changes reach a preset threshold. This step includes: continuously recording the market price change index of real estate within the geographic unit; when the market price change index of real estate within the geographic unit shows an upward trend and reaches a preset threshold, generating a price increase record for real estate within the geographic unit, which includes the time of the price increase; when the holding cost change index of real estate within the geographic unit shows an upward trend and reaches a preset threshold, retrieving a price increase record for real estate within the geographic unit within a preset time window; if a price increase record for real estate within the geographic unit exists within the preset time window, it is determined that the holding cost change index and market price change index of real estate within the geographic unit simultaneously show an upward trend within the same analysis period.
[0027] Among these, "showing an upward trend" refers to an indicator exhibiting a sustained growth trend over a period of time. This can be identified using time series analysis methods, such as moving averages, linear regression analysis, or trendline fitting techniques. Its purpose is to filter out short-term fluctuations and focus on long-term or medium-term data trends. "Preset threshold" refers to a pre-defined value or percentage used to measure whether the magnitude of an indicator's change reaches a significant level. It can be implemented using a fixed value, percentage increment, or a dynamic threshold based on historical data statistical distribution. Its purpose is to ensure that only significant changes with practical meaning are considered, avoiding misjudgments of minor fluctuations. "Price increase record" refers to a data entry used to store information related to when the market price change indicator of real estate within a geographic unit reaches the upward trend and preset threshold conditions. It can be implemented using a data structure including fields such as timestamp, geographic unit identifier, and price change magnitude. Its purpose is to provide historical evidence for subsequent cross-validation. "Preset time window" refers to a pre-defined time range used to limit the effective time interval when retrieving price increase records. It can be implemented using a fixed duration or a dynamically adjusted time range. Its purpose is to tolerate potential time lags between different indicators while avoiding the correlation of irrelevant events due to overly broad time ranges.
[0028] This application's solution addresses the lag and noise issues inherent in traditional methods when determining whether changes in holding costs and market prices of properties within a geographic unit simultaneously show an upward trend within the same analysis period. Specifically, the system first continuously monitors and records changes in the market price of properties within the geographic unit. When this indicator not only shows an upward trend but also reaches a preset threshold, the system immediately generates a record of the price increase within the geographic unit, explicitly including the specific time of the price increase. This step ensures that only significant market price increases are captured and timestamped, laying the foundation for subsequent cross-validation and avoiding misjudgments of short-term, insignificant fluctuations. Furthermore, when the system detects that the holding cost change indicator of properties within the geographic unit also shows an upward trend and its change also reaches a preset threshold, it does not immediately make a synchronous judgment. Instead, the system utilizes a preset time window to actively retrieve previously generated records of price increases within the geographic unit. This design cleverly considers the inherent time lag between different indicators. For example, an increase in holding costs may take some time to be transmitted and reflected in market prices. By searching within a certain time window, rather than requiring strict synchronization, the solution tolerates this natural time difference, thus avoiding underreporting due to data asynchrony. If a record of price increases for properties within the corresponding geographic unit is indeed found within this preset time window, the system can reliably determine that the holding cost change indicator and the market price change indicator for properties within the geographic unit are both showing an upward trend within the same analysis period. This judgment logic combines the significant upward trend of the two indicators with their temporal correlation, ensuring the accuracy and robustness of the judgment. This solution is closely integrated with the step in this application that determines whether the holding cost change indicator and the market price change indicator for properties within a geographic unit are both showing an upward trend within the same analysis period, and whether their respective changes reach preset thresholds. By providing a more accurate and robust judgment mechanism for "simultaneous upward trends," this solution significantly improves the accuracy of the entire automatic housing price valuation data anomaly detection method. It enables the system to more effectively identify complex anomaly patterns driven by the same event, but which may have temporal misalignment or noise interference in data performance, such as the contradictory phenomenon mentioned in the background technology, where the increase in property costs due to subway construction and the rise in the listing prices of surrounding properties increase.
[0029] In some preferred embodiments, this application is implemented as follows. To determine whether the holding cost change index and market price change index of properties within a geographic unit simultaneously show an upward trend within the same analysis period, and whether their respective changes reach preset thresholds, the system can continuously obtain market price data of properties within the geographic unit from the property transaction database and calculate its market price change index. For example, the system can update the data hourly or daily and use a seven-day moving average or linear regression slope to determine whether the market price change index shows an upward trend. When the moving average line rises for three consecutive days, and the current market price has increased by more than 5% (preset threshold) compared to thirty days ago, the system can generate a price increase record. This record can be a data entry containing the geographic unit ID, the timestamp of the price increase (e.g., "2023-10-26 10:30:00"), and the price change magnitude, and store it in memory or a temporary database. Simultaneously, the system can continuously monitor the holding cost change index of properties within the geographic unit, for example, by obtaining data such as property fees and taxes through a property management system. When a change in holding costs, such as monthly property fees, increases by more than 3% compared to the previous month (a preset threshold), and trend analysis (e.g., two consecutive months of month-on-month increases) indicates an upward trend, the system triggers a search operation. This search operation will look for records of price increases for properties within previously generated geographic units within a preset time window, such as the past 90 days. If the system successfully retrieves at least one record of price increases for that geographic unit within this 90-day time window, then the system can determine that both the change in holding costs and the change in market prices for properties within that geographic unit are showing an upward trend within the same analysis period.
[0030] This application further proposes steps for calculating the change index of holding costs of real estate within a geographic unit, including: identifying the management entity to which the real estate within the geographic unit belongs; aggregating the holding cost data of the management entity to obtain unified holding cost data at the management entity level; and using the unified holding cost data at the management entity level when calculating the change index of holding costs of real estate within the geographic unit.
[0031] Identifying the management entity to which a property belongs within a geographic unit refers to an organization or institution that has direct or indirect authority to manage, collect, or influence the holding costs of properties within that geographic unit. This can include property management companies, community management committees, owners' committees, relevant government departments, or other institutions responsible for maintaining and operating the common areas of the property. The purpose is to provide a logical basis for subsequent data aggregation, classifying scattered property data according to their respective management entities, thereby facilitating unified data processing. Specifically, the correspondence between properties and management entities can be obtained by querying property registration information, property service contracts, community management records, or relevant public databases. For example, matching and identification can be performed in a pre-established management entity database based on the property's address, community name, or property certificate number. Aggregating the holding cost data of management entities refers to summarizing, statistically analyzing, or calculating the holding cost data of multiple properties under the same management entity to form a single or set of data representing the overall holding cost level of that management entity. The purpose is to eliminate potential fluctuations or anomalies in individual property data, forming more representative and stable management entity-level data, thereby more accurately reflecting regional holding cost trends. Various statistical methods can be employed, such as calculating the average, weighted average, median, and sum, or establishing statistical models to comprehensively consider changes in various holding costs. Unified holding cost data at the management entity level refers to aggregated data that represents the overall holding cost level of a specific management entity within a certain time period or cycle. This data is no longer the original holding cost of a single property, but rather standardized and integrated data reflecting trends at the management entity level. Its purpose is to provide a standardized and consistent data foundation for calculating subsequent property holding cost change indicators within geographic units, ensuring the accuracy and comparability of the calculation results. Specifically, this can be represented as the aggregated average holding cost, the rate of change in holding costs, or the total holding cost at a specific point in time.
[0032] This application's solution addresses the accuracy issues that may arise when directly calculating holding cost change indicators from scattered property data within a geographic unit by introducing a management entity-level holding cost data processing mechanism. Specifically, the solution first identifies the management entity to which each property within the geographic unit belongs. This step is fundamental to data integration, logically categorizing property data that might otherwise be managed by different property companies or management agencies. Subsequently, the holding cost data of these identified management entities is aggregated to obtain unified holding cost data at the management entity level. This aggregation effectively smooths out noise and outliers in individual property data, forming a more representative data view reflecting the overall holding cost level of the management entity. Finally, when calculating the holding cost change indicator for properties within a geographic unit, the original, scattered property data is no longer used directly; instead, the aggregated unified holding cost data at the management entity level is employed. It is precisely this elevation of the hierarchy from individual properties to management entities and the unification of data processing that enables the calculated holding cost change indicator to more accurately and stably reflect the overall trend within the geographic unit. This approach avoids calculation biases caused by diverse data sources, inconsistent recording standards, or significant fluctuations in individual property data, thereby improving the accuracy and reliability of the holding cost change indicator. When this more accurate holding cost change indicator is used together with the market price change indicator to determine the degree of contradiction, it can more accurately identify the inherent logical contradictions between holding costs and market prices driven by the same external event. This provides a more solid data foundation for subsequent quantification of the degree of contradiction and identification of abnormal areas. This enables the entire automatic property valuation data anomaly detection method to more effectively discover and handle complex anomaly patterns that are difficult to identify using traditional methods, thus improving the overall accuracy and robustness of the valuation system.
[0033] In some preferred embodiments, this application is implemented as follows. Assume a geographical unit contains multiple residential areas, each managed by a different property management company. To calculate the change index of holding costs for properties within this geographical unit, the system first identifies the property management company for each property within the geographical unit. This can be achieved by querying property management company information recorded in a property database or by matching the property address with the property management company's service area. For example, for property A, the system identifies it as belonging to "Sunshine Property"; for property B, the system identifies it as belonging to "Harmony Property". Further, for each identified management entity, such as "Sunshine Property", the system collects holding cost data for all properties managed by it, including property management fees, public maintenance funds, parking fees, etc. Then, this data is aggregated. Specifically, the average change rate of property management fees and the average change in public maintenance funds for properties managed by "Sunshine Property" can be calculated to obtain unified holding cost change data at the "Sunshine Property" level. Similarly, similar aggregation processing is performed on the data of "Harmony Property" and all other management entities within the geographical unit to form unified holding cost data at their respective management entity levels. Finally, when calculating the change in holding costs of properties within the entire geographic unit, the system uses aggregated, unified holding cost data across all management entity levels. For example, the change in holding costs for a geographic unit can be obtained by weighted averaging of the unified holding cost data across all management entity levels, with the weights determined based on the number of properties managed or the total area managed by each entity. Therefore, the calculated change in holding costs for a geographic unit will be a comprehensive value reflecting the overall trend of management entities within the region, rather than simply a summary of data from all individual properties, thus avoiding calculation biases caused by individual differences or missing data.
[0034] This application further proposes steps for quantifying the degree of contradiction when the indicators of changes in holding costs and market prices of real estate within a geographic unit simultaneously show an upward trend within the same analysis period. These steps include: obtaining market activity information for the geographic unit; obtaining information on the types of real estate within the geographic unit; obtaining market trend information for the macroeconomic cycle; selecting contradiction quantification parameters suitable for the geographic unit from a preset parameter set based on the market activity information, the types of real estate within the geographic unit, and the market trend information for the macroeconomic cycle; and using the contradiction quantification parameters, calculating the degree of contradiction when the indicators of changes in holding costs and market prices of real estate within the geographic unit simultaneously show an upward trend within the same analysis period, based on the judgment results.
[0035] The preset parameter set refers to a series of selectable numerical or functional models prepared in advance to quantify the degree of contradiction in the real estate market. It can be implemented using a database containing different weighting factors, thresholds, nonlinear functions, or machine learning model parameters. The contradiction quantification parameter refers to a specific numerical value or function selected from the preset parameter set to specifically calculate the degree of contradiction when the indicators of changes in real estate holding costs and market price changes rise simultaneously within a geographical unit. It can be a weighting coefficient, an adjustment factor, or a specific calculation formula.
[0036] This application's approach, by introducing multi-dimensional contextual information, quantifies the degree of contradiction when indicators of changes in holding costs and market prices of properties within a geographic unit simultaneously show an upward trend within the same analysis period. Specifically, it first obtains market activity information for the geographic unit, reflecting the level of activity in real estate transactions within the region and providing a basis for assessing the market's tolerance for price fluctuations. Simultaneously, it obtains information on the type of properties within the geographic unit, as different types of properties exhibit varying sensitivities to market changes and costs, requiring differentiated treatment. Furthermore, it obtains market trend information related to the macroeconomic cycle, which helps understand the impact of the current overall economic environment on the real estate market, thereby enabling a more comprehensive assessment of the relationship between costs and prices.
[0037] After collecting this multi-dimensional information, the system intelligently selects the most suitable contradiction quantification parameters from a preset set of parameters, based on market activity information, property type information within the geographical unit, and macroeconomic cycle market trend information. This selection mechanism ensures the personalization and adaptability of the quantification process, avoiding biases that may result from using a single parameter. Subsequently, using the selected contradiction quantification parameters, combined with the previously determined upward trend of both holding cost change indicators and market price change indicators, the system accurately calculates the degree of contradiction within the properties of that geographical unit.
[0038] In this way, this solution elevates the step of quantifying the degree of contradiction from a simple, fixed-parameter calculation to a dynamic, context-aware computation. This allows subsequent steps of identifying and expanding anomalous regions based on the degree of contradiction to obtain more accurate input, thereby improving the accuracy and reliability of anomaly detection. This refined quantification enables the system to more accurately identify inherent logical contradictions between multi-feature data streams driven by the same event, thus effectively locating and addressing the resulting systemic and clustered valuation biases. This solves the problem that traditional methods struggle to accurately quantify the degree of contradiction in complex market environments.
[0039] In some preferred embodiments, when quantifying the degree of contradiction between changes in the holding cost of properties and changes in market prices within a geographic unit, both showing an upward trend within the same analysis period, the system first obtains market activity information for the geographic unit. This can be determined by analyzing data such as property transaction volume, number of listed properties, and average transaction cycle in the geographic unit over the past quarter. For example, if the transaction volume is higher than the historical average and the transaction cycle is shorter, it is judged as high market activity. Simultaneously, the system obtains information on the type of properties within the geographic unit. This can be identified by querying property registration data or planned use data, for example, distinguishing between ordinary residential properties, commercial apartments, and office buildings. Furthermore, the system also obtains market trend information related to the macroeconomic cycle. This can be determined by analyzing macroeconomic indicators such as the quarterly GDP growth rate released by the National Bureau of Statistics, benchmark interest rate adjustments announced by the central bank, and the real estate industry prosperity index to determine whether the current market is in an upward, downward, or stable phase.
[0040] After acquiring the above information, the system selects appropriate quantitative parameters for the current geographic unit from a pre-set set of parameters based on this multi-dimensional information. This pre-set set of parameters can be a structured database storing pre-configured quantitative models or weighting factors for different combinations of market activity, property type, and macroeconomic cycle. For example, if a geographic unit is identified as having "high" market activity, a "residential" property type, and an "upward" macroeconomic cycle, the system can automatically match and select a specific set of weighting parameters, such as a combination of parameters that shows higher tolerance for price increases and lower sensitivity to cost increases.
[0041] Finally, the system uses the selected contradiction quantification parameters, combined with the previously determined upward trend of both holding cost change indicators and market price change indicators, to calculate the degree of contradiction within a geographical unit. For example, if the selected contradiction quantification parameter is a non-linear function, this function can dynamically adjust the weights of the holding cost change indicators and market price change indicators based on market activity, property type, and macroeconomic cycle, thereby outputting a contradiction intensity value that more accurately reflects the actual situation. In this way, the quantification results can more accurately reflect the true contradiction situation under specific market conditions.
[0042] This application further proposes a step for identifying and expanding the overall anomalous region affected by contradictions based on the degree of contradiction, including: identifying geographical units whose contradiction degree reaches a preset contradiction threshold based on the degree of contradiction, and marking them as anomalous candidate units; obtaining the geographical attribute information of the anomalous candidate units; obtaining the contradiction pattern characteristics of the anomalous candidate units; merging non-adjacent anomalous candidate units on the geographical unit according to the geographical attribute information and contradiction pattern characteristics of the anomalous candidate units to form discrete anomalous regions; expanding adjacent anomalous candidate units on the geographical unit according to the geographical adjacency relationship of the anomalous candidate units to form continuous anomalous regions; and merging the discrete anomalous regions and continuous anomalous regions to form an overall anomalous region.
[0043] Geographic attribute information refers to non-transactional data describing the spatial location, environmental characteristics, and related infrastructure of a geographic unit. This can be achieved using geographic coordinates, administrative divisions, land use types, surrounding transportation networks, and the distribution of public service facilities. Its purpose is to provide spatial contextual information to facilitate subsequent spatial analysis and regional identification. Contradictory pattern characteristics refer to a set of attributes reflecting specific contradictory manifestations between indicators of changes in property holding costs and market price changes within a geographic unit. This can be achieved using information such as contradiction intensity, duration, scope of influence, and type. Its purpose is to characterize the inherent nature and patterns of contradictions, facilitating the precise classification and merging of anomalous areas. Geographic adjacency refers to the topological relationship between geographic units that are spatially adjacent or share boundaries. This can be determined using methods based on Voronoi diagrams, Delaunay triangulation, or shared boundary judgment. Its purpose is to identify spatially continuous anomalous areas, thereby reflecting the diffusion and aggregation of anomalous phenomena.
[0044] This application's solution overcomes the limitations of anomaly detection based solely on isolated geographical units by performing spatial analysis and pattern recognition on the quantified degree of contradiction among geographical units obtained in previous steps. Specifically, firstly, the system identifies geographical units whose degree of contradiction reaches a preset contradiction threshold based on the degree of contradiction of each geographical unit, and initially marks these units as anomaly candidate units. This screening process is the foundation for subsequent anomaly area identification, ensuring that contradictory units are considered. Subsequently, to understand the spatial distribution and intrinsic characteristics of these anomaly candidate units, the system further acquires their geographical attribute information and contradiction pattern characteristics. Geographical attribute information provides the spatial location and environmental background of the unit, while contradiction pattern characteristics reveal the specific manifestations of the contradiction between changes in holding costs and market price changes within it. Acquiring this information provides data support for subsequent area merging and expansion. Based on this, this solution combines two spatial aggregation strategies. On the one hand, based on the geographical attribute information and contradiction pattern characteristics of the anomaly candidate units, the system can merge those anomaly candidate units that are not geographically adjacent but share similar spatial backgrounds or contradictory manifestations, thereby forming discrete anomaly areas. This merging mechanism can capture anomalies that are discontinuously distributed but essentially homogeneous. On the other hand, considering the spatial diffusion of anomalies, the system expands geographically adjacent anomaly candidate units based on their geographical adjacency, forming continuous anomaly regions. This expansion mechanism ensures the identification of anomaly regions, enabling the identification of spatially spreading clusters of anomalies triggered by local events. Ultimately, by merging these discrete and continuous anomaly regions, this scheme forms a unified anomaly region. This integration method ensures that the identification of areas affected by contradictions considers both spatial continuity and discontinuous but pattern-similar anomalies, thereby enabling the location of systematic and clustered valuation bias areas.
[0045] In some preferred embodiments, this application is implemented as follows. Assuming a certain urban area, the degree of contradiction for each geographical unit has been calculated through previous steps. The system first identifies all geographical units with a contradiction degree exceeding a preset contradiction threshold based on these contradiction degrees, and marks these units as anomalous candidate units. For example, old residential unit A, located directly above a tunnel and affected by vibration, and residential units B and C, located around a subway station and whose asking prices have increased due to anticipated appreciation, may all be marked as anomalous candidate units. Subsequently, the system obtains the geographical attribute information of these anomalous candidate units. For example, unit A may be identified as "an old brick-concrete residential area located above an underground tunnel," while units B and C may be identified as "areas surrounding a newly built subway station, mainly commercial housing." Simultaneously, the system obtains their contradiction pattern characteristics. For example, the contradiction pattern of unit A may be "rising holding costs and fluctuating market prices," while the contradiction pattern of units B and C may be "rising market prices and stable holding costs." Based on this information, the system performs regional merging and expansion. Specifically, if there exists another geographical unit D, which is not geographically adjacent to unit A but is also an "old brick-concrete residential area located above another underground tunnel," and its contradictory pattern also manifests as "rising holding costs and fluctuating market prices," then units A and D can be merged to form a discrete anomaly region, reflecting similar anomalies caused by the underground project. Simultaneously, if units B and C are geographically adjacent and both exhibit a contradictory pattern of "rising market prices," then they can be expanded to form a continuous anomaly region, reflecting regional market anomalies brought about by the subway's benefits. Ultimately, the system merges these identified discrete and continuous anomaly regions to form a unified anomaly region. This unified anomaly region will include all geographical units affected by the subway construction event that exhibit contradictory characteristics, regardless of whether their cost increases are due to physical impacts or price increases due to planning benefits, and regardless of whether they are continuously or discretely distributed, thus comprehensively identifying the systemic and clustered valuation bias areas caused by this event.
[0046] This application further proposes a step for merging non-adjacent anomalous candidate units on a geographic unit to form a discrete anomalous region based on the geographic attribute information and contradictory pattern characteristics of the anomalous candidate units. The steps include: analyzing the geographic attribute information of the anomalous candidate units to identify units sharing geographic features; analyzing the contradictory pattern characteristics of the anomalous candidate units to identify units with similar contradictory patterns; and merging non-adjacent anomalous candidate units on a geographic unit based on units with shared geographic features and units with similar contradictory patterns to form a discrete anomalous region.
[0047] Among them, the geographic attribute information of the abnormal candidate unit refers to various data describing the geographical location of the abnormal candidate unit and its surrounding environment. It can be represented by data such as administrative division, transportation convenience, surrounding supporting facilities, topography, and land use type. Its purpose is to provide a basis for spatial correlation. The contradiction pattern characteristics of the abnormal candidate unit refer to the quantitative description of the opposition relationship between the data change trends within the abnormal candidate unit. It can be represented by data such as the increase, rate of change, and trend consistency of holding cost change indicators and market price change indicators. Its purpose is to reveal the internal logical relationship of data anomalies.
[0048] This application's solution achieves the merging of geographically non-adjacent anomalous candidate units through multi-dimensional analysis. Specifically, firstly, the geographical attribute information of the anomalous candidate units is analyzed to identify units sharing geographical features. This analytical step can initially screen anomalous candidate units that may be spatially related, even if they are not directly adjacent geographically, but may be affected by common factors due to being located in the same administrative region, having similar transportation networks, or sharing similar infrastructure. Subsequently, the contradictory pattern characteristics of the anomalous candidate units are analyzed to identify units exhibiting similar contradictory patterns. This analytical step focuses on data-level logic and can discover anomalous candidate units that, although not necessarily geographically adjacent, exhibit similar contradictory behaviors in holding costs and market price change trends. Finally, based on the aforementioned identified units sharing geographical features and units with similar contradictory patterns, the geographically non-adjacent anomalous candidate units are merged to form discrete anomalous regions. This method, which comprehensively considers geographical attributes and contradictory pattern characteristics, avoids judgment biases that may be caused by single-dimensional analysis, ensuring the effectiveness of merging and the rationality of regional division. In this way, even anomalous units that are not geographically adjacent but are affected by the same potential event can be identified and grouped into the same discrete anomalous region.
[0049] As a component of identifying and expanding overall anomalous regions affected by contradictions, this scheme provides a more refined mechanism for the formation of discrete anomalous regions. This allows subsequent expansion of continuous anomalous regions and merging of overall anomalous regions to be based on a more effective initial identification, thereby improving the ability of the entire automatic housing price valuation data anomaly detection method to identify complex anomaly patterns and locate valuation deviation areas.
[0050] In some preferred embodiments, this application is implemented as follows. Assume there are multiple anomalous candidate units that are not geographically adjacent but may all be affected by the same urban infrastructure construction event, such as the planning and construction of a subway line.
[0051] First, the geographical attributes of these candidate anomalous units can be analyzed. For example, for each candidate unit, information such as its administrative division, distance to the nearest subway station, density of the surrounding road network, and whether it is located within the subway construction impact zone can be extracted. By performing cluster analysis or rule matching on this geographical attribute information, units sharing geographical characteristics can be identified. For example, even if unit A and unit B are not geographically adjacent, if they are both located along the newly planned subway line, or are both in the same area and within a reasonable range of the planned subway station, they can be identified as units sharing the geographical characteristic of "subway planning impact".
[0052] Secondly, the contradictory patterns of these anomalous candidate units can be analyzed. For example, for each anomalous candidate unit, the specific values, magnitudes, rates of change, and trend consistency of its property holding cost change indicators (such as rising property management fees) and market price change indicators (such as rising listing prices) can be extracted. By performing pattern recognition or similarity calculations on these contradictory pattern characteristics, units with similar contradictory patterns can be identified. For example, if both Unit A and Unit B exhibit a contradictory pattern of a slight increase in property management fees due to construction impacts, while simultaneously experiencing a significant increase in listing prices due to the benefits of subway planning, they can be identified as units with similar contradictory patterns.
[0053] Finally, based on the identified units sharing geographical features and exhibiting similar contradictory patterns, non-adjacent anomalous candidate units can be merged to form discrete anomalous regions. Specifically, anomalous candidate units are merged into the same discrete anomalous region only when they share specific geographical features (e.g., all affected by subway planning) and exhibit similar contradictory patterns (e.g., all showing a contradiction between rising costs and rising prices). For example, if units A, B, and C are not geographically adjacent but are all located along the subway line and all exhibit a contradictory pattern of "increased construction costs and rising market prices," then they will be merged into a single discrete anomalous region. This method ensures the inherent consistency of the merged regions, clarifies the causes of anomalies, and thus improves the accuracy of anomalous region identification.
[0054] This application further proposes steps for adjusting the property valuation process based on anomaly markers, including: selecting a valuation sub-model corresponding to the anomaly marker; and using the valuation sub-model to perform valuation processing on the property data within the overall anomaly area, thereby adjusting the property valuation process.
[0055] Among them, anomaly labeling refers to the identification information attached to the real estate data within the identified overall anomaly area to characterize its anomaly pattern or anomaly characteristics. It can be implemented in the form of labels, codes, or data fields, and its purpose is to provide targeted processing basis for subsequent valuation processes. Valuation sub-model refers to a valuation model with specific valuation logic or algorithm that is trained or configured for specific types or patterns of real estate data anomalies. It can be implemented using rule-based models, statistical models, or machine learning models, and its purpose is to more accurately handle abnormal data that is difficult for conventional valuation models to handle effectively.
[0056] This application's solution addresses the inaccuracy of traditional valuation processes when handling complex anomaly data by introducing targeted valuation sub-models. Specifically, after identifying and marking anomalies in the property data within the overall anomaly area, this solution first intelligently selects a valuation sub-model highly corresponding to the specific anomaly characteristics implied by these markings. This selection process is based on a deep understanding of the processing capabilities of different anomaly patterns and valuation models, ensuring that for each unique anomaly, the most suitable professional model for handling that type of data is matched. This precise matching avoids the biases and limitations that may result from using a single, general-purpose valuation model. Subsequently, the selected valuation sub-model is used to perform refined valuation processing on the property data within the overall anomaly area. This approach fully leverages the advantages of the valuation sub-model in handling specific anomalies, effectively correcting valuation biases caused by data anomalies, thereby achieving dynamic and adaptive adjustments to the property valuation process. Furthermore, this solution is closely integrated with the previous steps of identifying and expanding the overall anomaly area affected by contradictions and marking anomalies in the property data. The previous steps quantified the degree of contradiction when both the holding cost change indicator and the market price change indicator showed an upward trend within the same analysis period, and identified the overall abnormal area affected by the contradiction, providing a clear indication of the scope and type of abnormal data for this solution. Building on this, this solution goes beyond mere anomaly identification and further provides a targeted valuation processing mechanism. This integration creates a closed loop for the entire automatic housing price valuation data anomaly detection method: from anomaly discovery, location, and labeling to final valuation process adjustment and correction. In this way, the system can effectively handle systematic and clustered valuation biases caused by inherent logical contradictions between multi-feature data streams driven by the same event, significantly improving the accuracy and reliability of property valuations and making the valuation results closer to actual market conditions.
[0057] In some preferred embodiments, this application is implemented as follows: Assume that in a certain urban area, due to subway construction, some older residential communities experience increased property management costs due to construction impacts, while newly built communities in the surrounding area see increased market listing prices due to the anticipated subway opening. After identifying these affected overall abnormal areas, the system will attach corresponding abnormal tags to the property data within the area. For example, for properties in older residential communities where costs have increased due to construction impacts, the abnormal tags may include features such as "construction impact" and "cost increase"; while for properties in newly built residential communities where prices have increased due to the anticipated subway opening, the abnormal tags may include features such as "subway benefit" and "market expectation". When adjusting the property valuation process based on the abnormal tags, the system first parses these abnormal tags. For example, when the abnormal tags are identified to include the features of "construction impact" and "cost increase", the system will select a "cost correction valuation sub-model" specifically designed to handle cost-increase anomalies from a pre-set valuation sub-model library. This sub-model may quantify the cost increase caused by specific construction impacts based on historical data and reflect it in the property valuation. Meanwhile, when anomaly markers are identified that contain the characteristics of "subway benefits" and "market expectations," the system selects a "market expectation valuation sub-model." This sub-model may combine factors such as subway line planning, station distance, and surrounding amenities to assess the future appreciation potential of the property and incorporate it into the valuation considerations. Subsequently, the system uses the selected valuation sub-model to process the valuation of the marked property data within the overall anomaly area. For example, the "cost-adjusted valuation sub-model" will value the property data of older communities, making appropriate adjustments based on the increase in costs on the basis of the conventional valuation; while the "market expectation valuation sub-model" will value the property data of newly built communities, making appropriate increases based on market expectations on the basis of the conventional valuation. In this way, the most suitable valuation logic can be adopted for different types of anomalies, thereby achieving a refined adjustment of the property valuation process and avoiding the simplistic treatment of all anomaly data in a one-size-fits-all manner.
[0058] This application further proposes a step of selecting an evaluation sub-model corresponding to an anomaly label based on the anomaly label, including: parsing the anomaly label to obtain the anomaly features contained in the anomaly label; identifying multiple evaluation sub-models that can handle the anomaly features based on the anomaly features; ranking the multiple evaluation sub-models according to the degree of adaptation of each evaluation sub-model to the anomaly features; and selecting the evaluation sub-model with the highest priority in the ranking results as the evaluation sub-model corresponding to the anomaly label.
[0059] Among them, anomaly tags refer to data labels that identify anomalies detected in real estate data. They can be represented in the form of strings, enumeration values, or structured data, and their purpose is to clearly indicate the type and scope of the anomalies. Anomaly features refer to the key attributes or information extracted from the anomaly tags that describe specific anomalies. They can be obtained by parsing the fields, values, or combinations thereof of the anomaly tags. For example, the feature "foundation settlement" can be extracted from the "foundation settlement anomaly" tag. Its purpose is to provide a specific processing basis for the subsequent selection of valuation sub-models. Valuation sub-models refer to specialized algorithms or models for valuing specific types of real estate data or specific market conditions. They can be implemented through machine learning models, statistical regression models, or expert rule systems, and their purpose is to provide diversified valuation capabilities to cope with valuation scenarios of varying complexity. Adaptability refers to the applicability or effectiveness of a valuation sub-model in handling specific anomaly features. It can be quantified by the model's performance on historical data, the model's coverage of the feature, or the matching degree between the model's internal logic and the feature. Its purpose is to provide a quantitative basis for selecting the optimal valuation sub-model.
[0060] This application's solution optimizes the process of selecting valuation sub-models based on anomaly markers through a series of logically progressive steps, thereby improving the accuracy of property valuation. First, after anomaly markers are attached to property data, the system parses these markers to extract the specific anomaly features they contain. This parsing process forms the basis for subsequent precise selection, transforming abstract anomaly markers into concrete problem descriptions that can be identified and processed by the models. For example, a "structural anomaly" marker might be parsed to reveal specific anomaly features such as "foundation settlement" or "wall cracks." Second, based on these explicit anomaly features, the system identifies all valuation sub-models capable of handling these features. This step expands the potential selection range, ensuring that no potentially applicable specialized models are overlooked. Different valuation sub-models may excel at handling different types of anomalies; for example, one sub-model might focus on structural problems, while another excels at handling market sentiment fluctuations. Based on this, the system ranks these models according to their suitability for the current anomaly features. This ranking process is the core of the solution; it quantifies the merits of each model in handling specific anomalies, ensuring the scientific rigor of the selection. The suitability assessment can be based on factors such as the model's past success rate in handling similar anomalies and the degree of matching between the model's area of expertise and the current anomaly characteristics. Finally, the system selects the highest-priority valuation sub-model from the ranking results and uses it as the valuation sub-model corresponding to the current anomaly label. This ensures that the selected valuation sub-model is best suited to handle the current complex anomaly, avoiding suboptimal selections that might result from simple matching. This refined valuation sub-model selection mechanism allows the most suitable professional model to be used to value property data within the overall anomaly area when adjusting the property valuation process based on anomaly labels. This contrasts sharply with simply selecting a preset model based on anomaly labels, significantly improving the targeting and accuracy of the valuation. For example, when property data is marked as anomaly due to foundation settlement, this solution can identify and prioritize a valuation sub-model that excels at handling structural problems, rather than a general model that might focus more on market supply and demand analysis. This ensures that the final property valuation result more accurately reflects the true value of the property, effectively avoiding valuation biases caused by inappropriate model selection.
[0061] In some preferred embodiments, this application is implemented as follows. Assume the system detects a composite anomaly in property data involving "foundation settlement" and "increased property management fees," and adds an anomaly marker containing this information to the property data. First, the system can parse the anomaly marker to obtain specific anomaly characteristics, such as "foundation settlement" and "increased property management fees." Next, the system can identify multiple valuation sub-models capable of handling these characteristics. For example, the system can identify "Structural Impact Valuation Model A" (adept at handling structural problems such as foundation settlement), "Cost Change Valuation Model B" (adept at handling changes in holding costs such as property management fees and taxes), and "Comprehensive Anomaly Valuation Model C" (capable of considering both structural and cost factors). Then, the system can rank these models based on their suitability for the two anomaly characteristics of "foundation settlement" and "increased property management fees." For example, by evaluating the historical performance of Model A in handling foundation settlement cases, the historical performance of Model B in handling property management fee increase cases, and the performance of Model C in handling composite anomalies, their suitability can be quantified. Assuming the evaluation results show that "Comprehensive Anomaly Valuation Model C" best fits the current complex anomaly, followed by Model A, and then Model B, the system can ultimately select "Comprehensive Anomaly Valuation Model C," with the highest priority, as the valuation sub-model corresponding to this anomaly marker. In subsequent valuation processing, the system will use the model best suited to handle the complex anomalies of foundation settlement and rising property fees, thus obtaining a more accurate valuation result.
[0062] This application further proposes a step for ranking multiple valuation sub-models based on the degree of adaptation of each valuation sub-model to the abnormal features, including: obtaining the abnormal features; establishing a historical database to record the historical performance data of each valuation sub-model when handling different abnormal features; retrieving historical performance data of valuation sub-models related to the abnormal features from the historical database based on the abnormal features; evaluating the degree of adaptation of each valuation sub-model to the abnormal features based on the retrieval results; and ranking the multiple valuation sub-models based on the evaluation results.
[0063] Among these, "abnormal features" refer to the specific attributes or patterns parsed from anomaly markers that describe abnormal situations in real estate data. Examples include abnormally simultaneous increases in housing prices and holding costs in a specific area, or abnormal fluctuations in transaction volume for a specific property type. These can be represented using structured data fields, text descriptions, or feature vectors. "Historical database" refers to a database or dataset used to store and manage past operational data of the valuation sub-model. It can be constructed using relational databases, NoSQL databases, or distributed file systems. "Historical performance data" refers to the performance indicators and results records generated by the valuation sub-model when handling specific abnormal features. Examples include valuation accuracy, valuation deviation, processing time, model convergence, and error rate or correction effects under specific abnormal scenarios. This data can be recorded using structured tables, log files, or serialized objects. "Adaptability" refers to the applicability and effectiveness of the valuation sub-model in handling specific abnormal features. It can be quantified using comprehensive scores, rankings, probability values, or error indicators, aiming to measure the performance of the valuation sub-model when facing specific abnormal features.
[0064] By identifying the anomaly characteristics that need to be processed, the target of model selection is clearly defined. Subsequently, a historical database is established and maintained, continuously recording the actual performance data of each valuation sub-model when handling different types of anomalies. This historical performance data provides an objective and quantitative basis for subsequent model evaluation. When a new anomaly appears, the system can accurately retrieve the past performance records of the valuation sub-models related to that feature from the historical database. Based on these retrieval results, the system can evaluate the suitability of each valuation sub-model for handling the current anomaly. This evaluation is based on the model's performance in real-world scenarios, rather than simple preset rules, thus ensuring the reliability of the evaluation results. Finally, based on the evaluation results, multiple valuation sub-models are ranked, allowing the model with the highest suitability to be selected first. It is precisely because of this dynamic evaluation and ranking mechanism based on historical performance data that the proposed solution ensures that the selected valuation sub-model is highly matched to the specific anomaly when adjusting the property valuation process.
[0065] In some preferred embodiments, when the system identifies the need to appraise property data within an overall abnormal area, it first obtains the anomalous features contained in the anomalous markers attached to the current property data. For example, if the anomalous markers indicate a combination of "abnormally rising holding costs" and "abnormally rising market prices" in the area, these specific details are extracted as anomalous features. Next, the system establishes a historical database, which can be a distributed database storing a large amount of historical performance data of valuation sub-models when handling various anomalous features. For example, the historical database may record that valuation sub-model A achieved a valuation accuracy of 98% and a valuation deviation of 0.5% when handling the "abnormally rising market prices" feature; and valuation sub-model B achieved an accuracy of 97% and a deviation of 0.8% when handling the "abnormally rising market prices" feature. After obtaining specific anomalous features, the system retrieves the historical performance data of valuation sub-models related to these features from the historical database. For example, if the current anomalous feature is "abnormally rising holding costs," the system queries the historical database to find all valuation sub-models that have handled this feature and their corresponding historical performance data. Based on retrieved historical performance data, the system evaluates the suitability of each valuation sub-model for the current anomalous features. The evaluation can employ a weighted scoring mechanism, for example, assigning higher weight to valuation accuracy and lower weight to valuation bias, to calculate a comprehensive suitability score. For instance, valuation sub-model A might have a comprehensive suitability score of 0.95 when handling the current anomalous features, while valuation sub-model C might have a comprehensive suitability score of 0.88. Finally, based on these evaluation results, the system ranks the multiple valuation sub-models, placing the one with the highest suitability score at the top, thus providing the optimal selection for subsequent valuation processing.
[0066] refer to Figure 2This application further proposes an automatic housing price valuation data anomaly detection system, applied to an automatic housing price valuation data anomaly detection method. The system includes: a geographic unit processing module, which divides the urban geographic space into multiple geographic units and aggregates property data to the corresponding geographic units based on the property's geographical location; an indicator calculation module, which calculates the holding cost change indicator and market price change indicator for each geographic unit; a judgment module, used to determine whether the holding cost change indicator and market price change indicator for properties within a geographic unit simultaneously show an upward trend within the same analysis period, and whether their respective changes reach preset thresholds; a quantification module, based on the judgment results, quantifying the degree of contradiction when the holding cost change indicator and market price change indicator for properties within a geographic unit simultaneously show an upward trend within the same analysis period; an anomaly area processing module, which identifies and expands the overall anomaly area affected by the contradiction based on the degree of contradiction; and an adjustment module, used to attach anomaly markers to the property data within the overall anomaly area and adjust the property valuation process based on the anomaly markers.
[0067] The geographic unit processing module is responsible for logically dividing the urban geographic space and classifying and organizing the raw real estate data according to its geographical location. Specifically, it can be a data preprocessing service or a Geographic Information System (GIS) component, aiming to lay the foundation for subsequent regional data analysis. The indicator calculation module is responsible for extracting and calculating key economic indicators reflecting the dynamics of the real estate market for each divided geographic unit. Specifically, it can be a data analysis engine or a statistical calculation service, aiming to provide quantitative evidence for identifying potential market contradictions. The judgment module is responsible for logically judging the calculated real estate holding cost change indicators and market price change indicators according to preset conditions and rules. Specifically, it can be a rule engine or a condition judge, aiming to initially screen out potentially abnormal indicators. The system is divided into three modules: a geographic unit for identifying contradictions; a quantification module, which assesses and quantifies the degree of potential contradictions based on the judgment results (such as an algorithm model or a scoring system), and an anomaly area processing module, which identifies and expands the affected geographic areas based on the quantified degree of contradictions (such as a spatial analysis algorithm or a cluster analyzer), and an adjustment module, which marks the affected property data based on the output of the anomaly area processing module and adjusts the subsequent property valuation process accordingly (such as a data marking service or a process control component), and an adjustment module, which ensures that abnormal data is properly handled and improves the accuracy of the valuation results.
[0068] Specifically, the geographic unit processing module first finely divides the urban geographic space and accurately aggregates massive amounts of real estate data into corresponding geographic units based on their geographical location. This foundational work ensures that all subsequent analyses can focus on spatially correlated sets of properties, providing a structured data foundation for regional analysis. Next, the indicator calculation module calculates, in parallel, the property holding cost change indicator and market price change indicator for each geographic unit. The calculation of these two indicators is crucial for quantifying market dynamics; they reflect two important dimensions of property value and provide necessary data support for subsequent conflict analysis. Subsequently, the judgment module receives and processes these indicators. It intelligently determines whether the property holding cost change indicator and market price change indicator simultaneously show an upward trend within the same analysis period, and whether their respective changes reach preset thresholds. This judgment mechanism is the core of identifying potential conflict areas; it effectively filters out normal market fluctuations and focuses attention on significant data patterns that may have inherent conflicts. Once this dual upward trend is identified, the quantification module intervenes. Based on the assessment results, it precisely quantifies the degree of contradiction when the indicators of changes in holding costs and market prices of properties within a geographical unit simultaneously show an upward trend within the same analysis period. This quantification process allows for a numerical assessment of the severity of the anomaly, providing a reliable basis for subsequent anomaly area identification. Building on this, the anomaly area processing module utilizes the quantified degree of contradiction to not only identify geographical units where the degree of contradiction reaches a preset threshold but also further considers spatial correlation, expanding the overall anomaly area affected by the contradiction. This means the system can pinpoint the core area of the problem and comprehensively present its impact, avoiding omissions that might occur with isolated analysis. Finally, the adjustment module adds anomaly tags to the property data within the identified overall anomaly area and intelligently adjusts the property valuation process based on these tags. In this way, the system can differentiate the processing of anomaly data; for example, it can trigger specific valuation sub-models or introduce a manual review mechanism, thereby effectively reducing the negative impact of anomaly data on the final valuation result and significantly improving the accuracy and reliability of property valuation. The entire system transforms the originally complex methods and steps into an automated and manageable process through close collaboration and data flow between modules. The geographic unit processing module provides the data foundation, the indicator calculation module provides the analytical data, the judgment module performs initial screening, the quantification module performs in-depth evaluation, the anomaly area processing module performs spatial positioning, and the adjustment module achieves the final intervention and optimization.
[0069] In some preferred embodiments, this application is implemented as follows. The geographic unit processing module can be deployed as a backend service based on a Geographic Information System (GIS). This service can access geographic boundary data provided by urban planning departments, such as administrative divisions, street grids, or custom hexagonal grids, and integrate a real estate transaction database. It can quickly and automatically aggregate newly collected real estate data (including latitude and longitude information) into the corresponding geographic units using spatial indexing technologies (such as R-trees or quadtrees). The indicator calculation module can be implemented as a computing node in a data processing pipeline, for example, using the Apache Spark or Hadoop MapReduce framework. It periodically reads historical transaction data and holding cost data for each geographic unit from the real estate database. Specifically, the holding cost change indicator can be calculated as the average growth rate of property fees, property taxes, loan interest, etc., within a specific time period, while the market price change indicator can be calculated as the increase in the average listing price or transaction price of properties within that geographic unit within the same time period. The judgment module can be configured as a decision service based on a rule engine, for example, implemented using Drools or a custom scripting language. This service receives the holding cost change indicator and market price change indicator output by the indicator calculation module and makes judgments based on preset logical rules. For example, the rule can be set as follows: if the average monthly growth rate of the holding cost change indicator exceeds 1% for two consecutive analysis periods (e.g., two months), and the average monthly growth rate of the market price change indicator also exceeds 2% within the same two months, then it is determined that they are showing an upward trend simultaneously and have reached a preset threshold. The quantification module can be a service based on a machine learning model, such as a support vector machine (SVM) or neural network model. This model learns historical cases of contradictions when holding costs and market prices rise simultaneously under different market activity levels, property types, and macroeconomic cycles. When the judgment module outputs that there is a simultaneous upward trend, the quantification module will input the market activity of the current geographical unit (e.g., transaction volume, listing volume), property type distribution (e.g., residential, commercial, apartment ratio), and macroeconomic trends (e.g., interest rates, GDP growth rate), and output a contradiction score between 0 and 1, with a higher score indicating a more significant contradiction. The anomaly region processing module can be a spatial analysis and clustering service, for example, implemented using Python's GeoPandas library combined with Scikit-learn's DBSCAN clustering algorithm. This service first identifies geographic units with a conflict score exceeding a preset threshold (e.g., 0.7) as anomaly candidates. Then, it analyzes the geographic adjacency of these candidate units, merging adjacent anomaly candidates to form contiguous anomaly regions. Simultaneously, it also analyzes non-adjacent anomaly candidates with similar conflict patterns (e.g., similar conflict quantification scores, similar property types), grouping them into discrete anomaly regions.Ultimately, these continuous and discrete regions are merged to form a complete overall anomaly region. The adjustment module can be a data management and process scheduling service. Once the anomaly region processing module identifies the overall anomaly region, the adjustment module automatically adds a structured anomaly marker, such as a JSON field, to all affected property data records within that region, containing information such as the anomaly type (e.g., "cost-price contradiction"), the degree of contradiction, and the scope of impact. Subsequently, the module triggers or notifies the downstream property valuation system to automatically select or suggest the use of specific valuation sub-models based on these anomaly markers (e.g., a depreciation model for areas with foundation settlement, or an expected value-added model for areas with planning benefits), or submits the data marked as anomaly to human experts for review, thereby ensuring that the valuation process can address anomalies in a targeted manner.
[0070] The content disclosed above is only a preferred and feasible embodiment of the present invention, and is not intended to limit the scope of protection of the present invention. Therefore, all equivalent technical changes made based on the content of the present invention specification and drawings are included within the scope of protection of the present invention. Furthermore, the elements therein can be updated as technology develops.
Claims
1. A method for detecting anomalies in automatic housing price valuation data, characterized in that, The method includes the following steps: The city's geographic space is divided into multiple geographic units, and the property data is collected into the corresponding geographic units according to the geographical location of the properties. For each geographic unit, calculate the corresponding indicators of changes in holding costs and market prices of properties within that geographic unit. Determine whether the indicators of changes in holding costs and market prices of real estate within a geographical unit show an upward trend simultaneously within the same analysis period, and whether the magnitude of each change reaches a preset threshold. Based on the judgment results, the degree of contradiction between the indicators of changes in the holding cost of real estate and changes in market price within the same geographical unit when they both show an upward trend in the same analysis period is determined. Based on the degree of contradiction, identify and expand the overall anomalous region affected by the contradiction; Add anomaly markers to property data within the overall abnormal area, and adjust the property valuation process based on the anomaly markers.
2. The method for detecting anomalies in automatic house price valuation data as described in claim 1, characterized in that, The steps to determine whether the indicators of changes in holding costs and market prices of real estate within a geographical unit show an upward trend simultaneously within the same analysis period, and whether the magnitude of each change reaches a preset threshold, include: Continuously record the market price change index of real estate within a geographic unit. When the market price change index of real estate within a geographic unit shows an upward trend and reaches a preset threshold, generate a price increase record of real estate within the geographic unit. The price increase record of real estate within a geographic unit includes the time when the price increase occurred. When the indicator of changes in the holding cost of real estate within a geographic unit shows an upward trend and reaches a preset threshold, the record of price increases of real estate within the geographic unit is retrieved within a preset time window. If there are records of rising property prices within a geographic unit within a preset time window, then it is determined that the indicators of changes in holding costs and market prices of properties within the geographic unit show an upward trend simultaneously within the same analysis period.
3. The method for detecting anomalies in automatic house price valuation data as described in claim 1, characterized in that, The steps for calculating the change in holding costs of real estate within a geographic unit include: Identify the management entity to which a property belongs within a geographic unit; Aggregate the holding cost data of the managed entities to obtain unified holding cost data at the management entity level; When calculating the change in holding costs of properties within a geographic unit, uniform holding cost data at the management entity level is used.
4. The method for detecting anomalies in automatic house price valuation data as described in claim 1, characterized in that, Based on the assessment results, the steps to quantify the degree of contradiction when the indicators of changes in holding costs and market prices of real estate within a geographical unit both show an upward trend within the same analysis period include: Obtain market activity information for geographic units; Obtain information on the type of property within a geographic unit; Obtain market trend information for the macroeconomic cycle; Based on market activity information of geographical units, real estate type information within geographical units, and market trend information of macroeconomic cycles, select contradiction quantification parameters suitable for geographical units from a preset parameter set; Using contradiction quantification parameters, based on the judgment results, the degree of contradiction is calculated when the indicators of changes in the holding cost of real estate and changes in market price within the same geographical unit show an upward trend simultaneously within the same analysis period.
5. The method for detecting anomalies in automatic house price valuation data as described in claim 1, characterized in that, Based on the degree of contradiction, the steps to identify and expand the overall anomalous region affected by the contradiction include: Based on the degree of contradiction, geographical units that reach a preset contradiction threshold are identified and marked as abnormal candidate units. Obtain the geographic attribute information of abnormal candidate units; Obtain contradictory pattern features of anomalous candidate units; Based on the geographical attribute information and contradictory pattern characteristics of the anomalous candidate units, non-adjacent anomalous candidate units on the geographical unit are merged to form discrete anomalous regions. Based on the geographical adjacency of the candidate anomalous units, the adjacent candidate anomalous units on the geographical unit are expanded to form a continuous anomalous region. The discrete and continuous abnormal regions are merged to form a unified abnormal region.
6. The method for detecting anomalies in automatic house price valuation data as described in claim 5, characterized in that, Based on the geographical attribute information and contradictory pattern characteristics of anomalous candidate units, the steps to merge non-adjacent anomalous candidate units to form discrete anomalous regions include: Analyze the geographic attribute information of anomalous candidate units to identify units that share geographic features; Analyze the contradictory pattern characteristics of abnormal candidate units and identify units with similar contradictory patterns; Based on units with shared geographical features and units with similar contradictory patterns, non-adjacent anomalous candidate units are merged to form discrete anomalous regions.
7. The method for detecting anomalies in automatic house price valuation data as described in claim 1, characterized in that, The steps involved in adjusting property valuations based on anomaly markers include: Based on the anomaly markers, select the valuation sub-model corresponding to the anomaly markers; By using a valuation sub-model, the property data in the overall abnormal area is valued, thereby adjusting the property valuation process.
8. The method for detecting anomalies in automatic house price valuation data as described in claim 7, characterized in that, The steps for selecting the valuation sub-model corresponding to the anomaly marker include: Parse the anomaly markers to obtain the anomaly characteristics contained within them; Based on the abnormal characteristics, identify multiple evaluation sub-models that can handle the abnormal characteristics; Based on the degree of adaptation of each valuation sub-model to the abnormal features, the multiple valuation sub-models are ranked. The evaluation sub-model with the highest priority in the ranking results is selected as the evaluation sub-model corresponding to the anomaly label.
9. The method for detecting anomalies in automatic house price valuation data as described in claim 8, characterized in that, The steps for ranking multiple valuation sub-models based on their fit to abnormal features include: Obtain abnormal characteristics; Establish a historical database to record the historical performance data of each valuation sub-model when dealing with different anomalies; Based on the abnormal characteristics, retrieve the historical performance data of the valuation sub-models related to the abnormal characteristics from the historical database; Based on the search results, the degree of fit of each valuation sub-model to the abnormal features is evaluated; Based on the evaluation results, the multiple valuation sub-models are ranked.
10. A system for detecting anomalies in automatic housing price valuation data, applied to the method for detecting anomalies in automatic housing price valuation data as described in claim 1, characterized in that, The system includes: The geographic unit processing module divides the urban geographic space into multiple geographic units and aggregates property data into the corresponding geographic units based on the geographical location of the properties. The indicator calculation module calculates the change in holding costs and market price of real estate within each geographic unit. The judgment module is used to determine whether the holding cost change index and market price change index of real estate within a geographical unit show an upward trend simultaneously within the same analysis period, and whether the change magnitude of each reaches a preset threshold. The quantitative module quantifies the degree of contradiction when the indicators of changes in holding costs and market prices of real estate within a geographical unit show an upward trend simultaneously within the same analysis period, based on the judgment results. The abnormal region processing module identifies and expands the overall abnormal region affected by the contradiction based on the degree of contradiction. The adjustment module is used to add anomaly markers to the property data within the overall abnormal area and adjust the property valuation process based on the anomaly markers.