Land evaluation method, device, equipment and medium
Patent Information
- Application Number
- CN202610827146.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-25
AI Technical Summary
[0002]目前,在城镇土地管理与存量盘活领域,对土地利用效能(如“低效利用”识别)的评价主要依赖以下几类方法:基于宏观统计数据的指标法、专家经验评价法、基础机器学习分类法;但现有方法对于所需评价的土地都采用统一评价标准体系,但不同类型的土地的价值程度所取决的因素不尽相同,因此导致评价不准确
[0006]根据本发明实施例的一种土地评价方法,至少具有如下有益效果:本发明首先处理原始土地数据,再根据不同区位处理原始土地特征得到相对区位特征,然后对相对区位特征进行筛选,确定原始土地数据对应的土地类别的专属特征子集,即得到用于训练各土地类型分别对应的类别提升评价模型的训练数据。本发明通过相对区位特征、各类别的类别提升评价模型将不同类型土地的评价指标分开,提高评价土地的准确度。
Smart Images

Figure CN122819946A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban planning, and more particularly to a land evaluation method, apparatus, equipment, and medium. Background Technology
[0002] Currently, in the field of urban land management and stock revitalization, the evaluation of land use efficiency (such as the identification of "inefficient use") mainly relies on the following methods: index method based on macro statistical data, expert experience evaluation method, and basic machine learning classification method. However, existing methods adopt a unified evaluation standard system for the land to be evaluated, but the factors that determine the value of different types of land are not the same, thus leading to inaccurate evaluation. Summary of the Invention
[0003] This invention aims to at least solve one of the technical problems existing in the prior art. To this end, this invention proposes a land evaluation method that separates the evaluation indicators for different types of land by improving the evaluation model through relative location characteristics and category-based evaluation, thereby improving the accuracy of land evaluation.
[0004] The present invention also proposes apparatus, equipment and media having the above-mentioned land evaluation method.
[0005] A land evaluation method according to a first aspect of the present invention includes: Obtain raw land data; The original land data is preprocessed to determine the relative location characteristics; By filtering the relative location features, a subset of land-specific features corresponding to the original land data is obtained, thereby obtaining a category improvement evaluation model for each land type.
[0006] A land evaluation method according to an embodiment of the present invention has at least the following beneficial effects: The present invention first processes the original land data, then processes the original land characteristics according to different locations to obtain relative location characteristics, and then filters the relative location characteristics to determine the exclusive feature subset of the land category corresponding to the original land data, that is, to obtain the training data used to train the category improvement evaluation model corresponding to each land type. The present invention separates the evaluation indicators of different types of land through relative location characteristics and category improvement evaluation models of each category, thereby improving the accuracy of land evaluation.
[0007] According to some embodiments of the present invention, the preprocessing of the original land data to determine relative location characteristics includes: The original land data is sequentially filtered for noise and redundancy features and cleaned to prevent cheating, resulting in training and validation data. The evaluation results data in the original land data are mapped to numerical values.
[0008] According to some embodiments of the present invention, the preprocessing of the original land data to determine relative location characteristics includes: Divide the training and validation data by the area of each plot to obtain the per-plot characteristics.
[0009] According to some embodiments of the present invention, the preprocessing of the original land data to determine relative location characteristics further includes: The ratio of the land area characteristic to the average land area characteristic of the same land type is taken as the overall relative characteristic; The location benchmark value is determined based on the average land-to-land characteristics and confidence weight of the same land type; The ratio of the average location feature to the location benchmark value is used as the relative location feature.
[0010] According to some embodiments of the present invention, the determination of the confidence weight includes: The confidence weight is the ratio between the number of plots in the location and the sum of the number of plots in the location and the smoothing coefficient.
[0011] According to some embodiments of the present invention, the step of filtering the relative location features to obtain a subset of land-type-specific features corresponding to the original land data, thereby obtaining a category improvement evaluation model corresponding to each land type, includes: Prefit the relative location features to obtain the information gain of each relative location feature; Based on the information gain, the relative location features are sorted to obtain a subset of land-specific features corresponding to the original land data; The specific feature subset is used to train and validate the corresponding land type category improvement evaluation model.
[0012] According to some embodiments of the present invention, training and validating the category enhancement evaluation model corresponding to the specific feature subset includes: If the category enhancement evaluation model fails the validation, then based on the information gain, features in the exclusive feature subset are removed and the removed features are disabled; based on the updated exclusive feature subset, the category enhancement evaluation model is retrained.
[0013] A land evaluation apparatus according to a second aspect of the present invention includes: The data collection module is used to acquire raw land data; The data preprocessing module is used to preprocess the original land data to determine the relative location characteristics; The model training module is used to filter the relative location features to obtain a subset of features specific to the land type corresponding to the original land data, thereby obtaining a category improvement evaluation model corresponding to each land type.
[0014] An electronic device according to a third aspect of the present invention includes: Memory, used to store programs; A processor for executing a program stored in the memory, wherein when the processor executes the program stored in the memory, the processor is configured to perform the method as described in any one of the first aspects.
[0015] According to a fourth aspect of the present invention, a storage medium stores computer-executable instructions for performing the method as described in any one of the first aspects.
[0016] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the description, claims, and drawings. Attached Figure Description
[0017] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation on the technical solutions of the present invention.
[0018] Figure 1 This is a flowchart of a land evaluation method provided in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0020] It should be understood that in the description of the embodiments of the present invention, "multiple" (or "amounts") means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. If "first," "second," etc., are used in the description, they are only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.
[0021] like Figure 1 As shown, an embodiment of the present invention provides a land evaluation method, including: Step S100: Obtain raw land data; Step S200: Preprocess the raw land data to determine the relative location characteristics; Step S300: Filter relative location features to obtain a subset of exclusive features for each land type corresponding to the original land data, thereby obtaining the category improvement evaluation model for each land type.
[0022] This invention first processes the raw land data, then processes the raw land characteristics according to different locations to obtain relative location features. Next, it filters these relative location features to determine the specific feature subsets for each land category corresponding to the raw land data. This results in training data used to train category improvement evaluation models for each land type. This invention separates the evaluation indicators for different land types through relative location features and category improvement evaluation models, thereby improving the accuracy of land evaluation.
[0023] In one embodiment, in step S100, the original land data includes: physical characteristics of land and buildings, including land type name and location; activity indicators include: transaction amount, unit turnover, tax revenue, registered capital, economic output, and nighttime light intensity; and human evaluation conclusions include: evaluation results, result text, total potential score, and potential classification.
[0024] In one embodiment, in step S200, the raw land data is preprocessed to determine relative location characteristics, including: The original land data is sequentially filtered for noise and redundancy features and cleaned to prevent cheating, resulting in training and validation data. The evaluation results data in the original land data are mapped to numerical values.
[0025] Before model training, in order to provide clean data for subsequent processing, static filtering rules based on statistical rules and expert knowledge are first established to remove explicit noise and logically invalid fields. The preprocessing process includes: Remove invalid and null columns: Automatically detect and delete fields with a high null value rate (such as "construction year") or fields containing only a single value; Remove irrelevant identifiers: delete database indexes and administrative codes that do not contribute substantially to the evaluation; Eliminating redundant hierarchical features: For situations where the same attribute has both "original values" and "manual hierarchical" (such as "tax" and "tax hierarchy"), the principle of "retaining continuous values and eliminating discrete hierarchicals" is adopted. Retaining original values (such as absolute tax values) can provide richer information density for the model, while eliminating manually preset hierarchical fields avoids information redundancy and interference from human thresholds.
[0026] Merging redundant semantics: For fields with highly redundant semantics, only fields with more standard classification systems are retained, reducing the dimensionality of the feature space; Anti-fraud cleaning: To ensure that the model derives its current efficiency based on the physical and economic characteristics of the land itself, rather than directly copying conclusions, a feature blacklist is constructed based on expert knowledge. Fields describing the future potential of the land parcel or those with strong subjective biases (such as "total potential score," "potential grade," "redevelopment status," etc.) are removed to prevent the introduction of human interference factors into the future potential score. Fields that directly contain evaluation conclusions are strictly removed, including "evaluation results" (scores) and "result text" (Labels / target columns) to prevent direct data leakage from the model. This process actually separates the evaluation results from the features upon which the evaluation results depend. Finally, the target column "Result Text" is read, and the text labels are mapped to ordered numerical values. Specifically, the model defines the mapping relationship as {'Inefficient utilization': 0, 'Medium utilization': 1, 'Intensive utilization': 2}. If the land parcel data is missing a label, it is automatically filtered out from the training set.
[0027] In one embodiment, in step S200, the raw land data is preprocessed to determine relative location characteristics, including: Divide the training and validation data by the area of each plot to obtain the per-plot characteristics.
[0028] Given the inherent differences in output efficiency among different land use types (such as industrial and commercial), and the significant impact of geographical environment on the economic level of different locations, it is difficult to objectively evaluate the true efficiency of a plot of land using only absolute values. Therefore, constructing the per-plot characteristics of each plot aims to eliminate background interference from type thresholds and spatial interests, and ensure the relative objectivity of the evaluation.
[0029] In one embodiment, step S200, which involves preprocessing the raw land data to determine relative location characteristics, further includes: The ratio of the average characteristic per unit area to the average characteristic per unit area of the same land type is taken as the overall relative characteristic; The location benchmark value is determined based on the average land-to-land characteristics and confidence weight of the same land type; The ratio of the average location feature to the location benchmark value is used as the relative location feature.
[0030] It is easy to understand that land includes several plots; the characteristic per unit area is a general term for the data.
[0031] In one embodiment, the confidence weight is determined by: The confidence weight is the ratio of the number of plots in the location to the sum of the number of plots in the location and the smoothing coefficient.
[0032] To eliminate the interference of plot size on total output and restore land use intensity, we construct the per-plot characteristic based on the plot's geometric area (Shape_Area); for the... Total indicators (Such as registered capital, economic output, nighttime light intensity, fixed asset investment, etc.), its per capita characteristics The calculation formula is:
[0033] The area of the land parcel. To prevent division by zero errors with extremely small values (e.g., 0.1), the code's preset indicator list is automatically traversed to generate new feature columns such as "Economic Output per Unit Area" and "Number of Enterprises per Unit Area". Given the inherent differences in output efficiency among different land use types (such as industrial and commercial), and the significant influence of geographical environment on the economic level of different locations, it is difficult to objectively evaluate the true efficiency of a plot using only absolute values. Therefore, we construct overall relative characteristics and locational relative characteristics to eliminate background interference from type thresholds and spatial interests, ensuring the relative objectivity of the evaluation; for the overall relative characteristics, we calculate the plot attribute values. Compared with the average of all similar land uses The ratio of these values is used to measure their relative level in the global dimension:
[0034] To address the issue of large biases arising from directly calculating the location mean due to insufficient sample sizes in some administrative or functional zones, we introduce Bayesian smoothing. We define the smoothed location benchmark value. for: ,in, This represents the sample mean within that location. The global mean; The confidence weight is calculated using the following formula: In the formula, This represents the number of samples within that location. The smoothing coefficient (in this embodiment) When the sample size When it is extremely small, As the baseline value approaches 0, it automatically converges to the global mean, avoiding feature distortion caused by fluctuations in small samples; ultimately, the corrected relative location features are obtained.
[0035] In one embodiment, in step S300, relative location features are filtered to obtain a subset of land-specific features corresponding to the original land data, thereby obtaining a category improvement evaluation model for each land type, including: Prefit the relative location features to obtain the information gain of each relative location feature; Based on information gain, the relative location features are sorted to obtain a subset of features specific to the land type corresponding to the original land data; The evaluation model is improved by training and validating a subset of specific features for the corresponding land type.
[0036] It is easy to understand that after the category improvement evaluation model is validated, the land data of the plots to be evaluated, i.e., the plot data, is input into the category improvement evaluation model, and the category improvement evaluation model outputs the evaluation results. The process of training and validating the model is similar to the existing training and validating model process, that is, the exclusive feature subset is divided into a training set and a validation set. The training set is used to train the model and adjust the model parameters, while the validation set is used to verify the accuracy of the model.
[0037] In one embodiment, training and validating a category improvement evaluation model for the corresponding land type using a subset of specific features includes: If the category enhancement evaluation model fails validation, then based on information gain, features in the exclusive feature subset are removed and the removed features are disabled; based on the updated exclusive feature subset, the category enhancement evaluation model is retrained.
[0038] Based on the aforementioned specific feature subset, divided into training and validation sets, land use efficiency is classified and predicted using the CatBoost ensemble learning algorithm (class enhancement algorithm), and a full-process diagnostic mechanism is introduced, specifically including the following steps: To address the issue that a fixed indicator system cannot adapt to different land use types, we adopt a dynamic screening strategy. First, we use a lightweight CatBoostClassifier (100 iterations) to prefit all features and extract feature importance. Based on information gain, we automatically extract the top 10 features to form a feature subset specific to the current land use type (such as industrial land or commercial land) for subsequent model training. To address hidden data leakage (such as a feature whose name is unrelated but whose value is highly correlated with the result), an automated closed loop of "detection-removal-retraining" is established: the filtered feature set is used for model trial training, and the accuracy of the test set is calculated in real time. Set an abnormally high accuracy threshold. (In this embodiment, it is set to 98%), if The system determines that a data leak has occurred. At this point, the system does not output a result, but instead uses the feature importance interface to reverse-engineer the top-3 features with the highest contribution, identifies them as "leak sources," automatically adds them to the blacklist, and permanently removes them from the feature pool. The system automatically triggers retraining, repeating the above steps until the test set accuracy falls back to the normal range. ), to ensure that the final model does not contain any logical cheating; After passing the leakage prevention test, the system enters the final verification phase, which focuses on evaluating the model's generalization ability: execution The validation process is repeated 10 times, each time using a different random seed for stratified sampling to ensure consistent proportions of each type of sample; the average training set accuracy is then summarized over 10 iterations. ) and average test set accuracy ( ), calculate generalization error Tiered response strategy: Healthy state: If the generalization error is within a reasonable range, output the final evaluation result and remediation suggestions; Overfitting risk: If (In this embodiment, the percentage is set to 15%). Instead of automatically deleting features, the system outputs an "overfitting risk warning report" to prompt manual intervention to check the data distribution or adjust the model depth to avoid accidentally deleting valid features.
[0039] The present invention also provides a land evaluation device, characterized in that it comprises: The data collection module is used to acquire raw land data; The data preprocessing module is used to preprocess the raw land data in order to determine the relative location characteristics; The model training module is used to filter relative location features to obtain a subset of features specific to each land type corresponding to the original land data, thereby obtaining a category improvement evaluation model for each land type.
[0040] It should be noted that the apparatus provided in this embodiment is used to perform any of the above-mentioned land evaluation methods.
[0041] This invention also provides an electronic device, which includes, but is not limited to: Memory, used to store programs; The processor is used to execute programs stored in memory. When the processor executes programs stored in memory, it is used to perform the land evaluation method described above.
[0042] The processor and memory can be connected via a bus or other means.
[0043] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the method described in the embodiments of the present invention. The processor implements the above method by running the non-transitory software program and instructions stored in the memory.
[0044] The memory may include a program storage area and a data storage area, wherein the program storage area may store the operating system and application programs required for at least one function; the data storage area may store data for executing the methods described above. Furthermore, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0045] The non-transitory software program and instructions required to implement the above terminal selection method are stored in memory and are executed by one or more processors.
[0046] This invention also provides a storage medium storing computer-executable instructions for performing the above-described methods.
[0047] In one embodiment, the storage medium stores computer-executable instructions that are executed by one or more control processors.
[0048] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0049] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0050] This document describes embodiments of the invention, including preferred embodiments known to the inventors for carrying out the invention. Variations of these embodiments will become apparent to those skilled in the art upon reading the foregoing description. The inventors encourage those skilled in the art to adopt such variations as appropriate, and the inventors intend to practice embodiments of the invention in ways other than those specifically described herein. Therefore, the scope of the invention includes all modifications and equivalents of the subject matter set forth in the appended claims, as permitted by applicable law. Furthermore, the scope of the invention covers any combination of the foregoing elements in all possible variations thereof, unless otherwise indicated herein or otherwise clearly contradicted by the context.
Claims
1. A land evaluation method, characterized in that, include: Obtain raw land data; The original land data is preprocessed to determine the relative location characteristics; By filtering the relative location features, a subset of land-specific features corresponding to the original land data is obtained, thereby obtaining a category improvement evaluation model for each land type.
2. The land evaluation method according to claim 1, characterized in that, The preprocessing of the original land data to determine relative location characteristics includes: The original land data is sequentially filtered for noise and redundancy features and cleaned to prevent cheating, resulting in training and validation data. The evaluation results data in the original land data are mapped to numerical values.
3. The land evaluation method according to claim 2, characterized in that, The preprocessing of the original land data to determine relative location characteristics includes: Divide the training and validation data by the area of each plot to obtain the per-plot characteristics.
4. The land evaluation method according to claim 3, characterized in that, The preprocessing of the original land data to determine relative location characteristics also includes: The ratio of the land area characteristic to the average land area characteristic of the same land type is taken as the overall relative characteristic; The location benchmark value is determined based on the average land-to-land characteristics and confidence weight of the same land type; The ratio of the average location feature to the location benchmark value is used as the relative location feature.
5. A land evaluation method according to claim 4, characterized in that, The confidence weights are determined in the following ways: The confidence weight is the ratio between the number of plots in the location and the sum of the number of plots in the location and the smoothing coefficient.
6. The land evaluation method according to claim 1, characterized in that, The process of filtering the relative location features yields a subset of land-specific features corresponding to the original land data, thereby creating a category improvement evaluation model for each land type, including: Prefit the relative location features to obtain the information gain of each relative location feature; Based on the information gain, the relative location features are sorted to obtain a subset of land-specific features corresponding to the original land data; The specific feature subset is used to train and validate the corresponding land type category improvement evaluation model.
7. A land evaluation method according to claim 6, characterized in that, The step of training and validating the category improvement evaluation model corresponding to the specific feature subset includes: If the category enhancement evaluation model fails the validation, then based on the information gain, features in the exclusive feature subset are removed and the removed features are disabled; based on the updated exclusive feature subset, the category enhancement evaluation model is retrained.
8. A land evaluation device, characterized in that, include: The data collection module is used to acquire raw land data; The data preprocessing module is used to preprocess the original land data to determine the relative location characteristics; The model training module is used to filter the relative location features to obtain a subset of features specific to the land type corresponding to the original land data, thereby obtaining a category improvement evaluation model corresponding to each land type.
9. An electronic device, characterized in that, include: Memory, used to store programs; A processor for executing a program stored in the memory, wherein when the processor executes the program stored in the memory, the processor is configured to perform the method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The device stores computer-executable instructions for performing the method as described in any one of claims 1 to 7.