Multisource data entropy correction G1 empowerment flood risk zoning method

By using the multi-source data entropy correction G1 weighting method, the problem of inconsistent standards in flood risk zoning of multi-source indicator data was solved, achieving stable reproduction and traceability of risk zoning results, and ensuring unified visualization of risk zoning maps and data reliability.

CN122045867APending Publication Date: 2026-05-15INFORMATION CENT OF YELLOW RIVER WATER RESOURCES COMMISSION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INFORMATION CENT OF YELLOW RIVER WATER RESOURCES COMMISSION
Filing Date
2026-02-06
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve quality control and standardization when multi-source indicator data exhibits spatiotemporal differences, omissions, and anomalies. This results in inconsistent risk zoning results that are difficult to trace, particularly in flood risk zoning applications where inconsistencies in the quality control indicator matrix and amplified fluctuations in difference measurements occur.

Method used

The G1 weighting method, which corrects entropy from multi-source data, involves obtaining zoning task identifiers, evaluation unit version identifiers, and data source version identifiers; performing data acquisition and metadata management to generate an original indicator set; performing quality control and caliber alignment to generate a quality control indicator matrix; calculating entropy values ​​and difference measures to generate an entropy difference set; generating adjacent importance ratios; performing G1 reverse recursion to generate a weight vector; performing weighted summation to generate a comprehensive risk score sequence; and performing K-means clustering to generate five-level risk labels and perform GIS-based visualization.

Benefits of technology

It achieves stable reproduction and traceability of risk zoning results under fluctuating conditions of multi-source indicator data. It ensures the consistency of the weight vector generation process by generating adjacent importance ratios through stable terms, monotonic constraints and upper and lower limit truncation rules, and supports unified visualization of five-level risk labels and risk zoning maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045867A_ABST
    Figure CN122045867A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing of disaster risk management and geographic information, in particular to a flood risk zoning method based on multi-source data entropy correction G1 empowerment. The method comprises the steps of obtaining multi-source data and performing metadata management to generate an original index set; performing quality control and aperture alignment to form a quality control index matrix; calculating entropy values based on the quality control data, and constructing an entropy difference set with difference measurement; generating an adjacent ratio mark set through stable item addition, approach threshold rollback, upper and lower limit truncation and monotone constraint; generating a weight vector by adopting hierarchical grouping and G1 inverted recursion; carrying out weighted summation to obtain a comprehensive risk scoring sequence; generating a five-level risk tag through K-means clustering; and finally, outputting a result data packet and a risk zoning map through GIS association visualization. According to the method, objectification, automation and traceability of weight distribution are realized through an entropy correction G1 weighting method, and the accuracy and practicability of flood risk zoning are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of disaster risk management and geographic information data processing, and in particular to a multi-source data entropy correction G1 weighted flood risk zoning method. Background Technology

[0002] In the field of disaster risk management and geographic information data processing, existing solutions typically revolve around a chain of steps including indicator system construction, data preprocessing, directional standardization, weighted calculation, comprehensive risk scoring, and hierarchical visualization. These solutions suffer from limitations such as reliance on empirically determined weights or difficulty in recalculation, sensitivity to fluctuations in the quality of multi-source indicator data, and difficulty in consistently reproducing hierarchical results. Existing methods often perform quality control and standardization under conditions of spatiotemporal differences, missing data, and anomalies in multi-source indicator data. However, they are constrained by the need for unified organization and traceable recording of criteria alignment, missing data gating, hierarchical imputation, anomaly pruning, and quality labeling. In scenarios involving multiple data updates or cross-regional assessment unit migrations, issues such as inconsistent quality control indicator matrix calibers and amplified fluctuations in difference measurements arise, making it difficult to achieve a stable comprehensive risk scoring sequence. For the joint processing of the weighting chain consisting of quality control index matrix, entropy difference set, adjacent importance ratio, adjacent ratio label set and G1 reverse recursion, existing technologies generally lack unified and operable constraints in the generation rules of adjacent importance ratio, auditable boundaries after the introduction of stable terms, triggering and recording of monotonic constraints and upper and lower limit truncation, and judgment and traceability of near threshold backoff. It is difficult to form a consistent process of collection-alignment-quality control-standardization-weighting-clustering-recording-GIS association visualization in the application scenario of flood risk zoning. This leads to inconsistencies and difficulty in traceability between risk zoning maps and result data packages during review, reproduction and continuous updating. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention provides a multi-source data entropy-corrected G1-weighted flood risk zoning method, comprising: S100: Obtain the zoning task identifier, evaluation unit version identifier and data source version identifier, perform data acquisition and metadata management processing, and generate the original indicator set; S200: Based on the original indicator set, perform quality control and caliber alignment processing to generate a quality control indicator matrix; S300. Based on the quality control index matrix, perform entropy value and difference measurement calculation processing to generate an entropy difference set; the entropy difference set includes evaluation unit identifier, index field, standardized value field, entropy value field and difference measurement field. S400. Based on the entropy difference set, perform adjacent importance ratio generation processing to generate an adjacent ratio label set; the adjacent ratio label set includes adjacent importance ratio fields associated with two adjacent indicator fields and generation labels; S500: Based on the adjacent ratio label set and entropy difference set, perform G1 reverse recursive generation of weight vector processing to generate weight vector; S600: Based on the weight vector and quality control index matrix, perform weighted summation to generate a comprehensive risk score sequence; S700: Based on the comprehensive risk score sequence, perform K-means clustering to generate five-level risk labels, and generate five-level risk labels and parameter records; S800, based on five-level risk labels and parameter records, performs GIS-related visualization to generate result data packages, and generates result data packages and risk zoning maps.

[0004] Furthermore, the process of performing data acquisition and metadata management includes: The evaluation unit boundary data undergoes uniqueness verification and topology consistency checks, including surface feature closure checks, overlap and conflict checks, and hole anomaly checks. Multi-source indicator data undergoes access normalization processing, including data source connection, file parsing, field extraction, and structured encapsulation. The multi-source indicator data includes tabular, raster, and time-series indicator data. Furthermore, spatiotemporal caliber metadata is extracted from the data source header information, product description fields, and caliber dictionary, and the unit caliber identifier undergoes legality verification. A correlation index is established between the spatiotemporal caliber metadata and the multi-source indicator data.

[0005] Furthermore, the original set of indicators includes: The original indicator set includes evaluation unit identifier field, boundary geometry field, indicator field, value field, time stamp field, spatial positioning field, data source identifier field, version field, and caliber field.

[0006] Furthermore, the process of performing quality control and caliber alignment includes: Spatiotemporal caliber alignment involves unifying the time stamp field, spatial location field, and caliber field under the constraints of the target time granularity identifier and the target spatial caliber identifier. This spatiotemporal caliber alignment includes time granularity unification, spatial caliber alignment, and unit caliber conversion. Missing data gating determines the missing status and proportion of value fields and triggers hierarchical imputation or deletion actions. Hierarchical imputation generates replacement values ​​for missing items according to three rules: priority of same source and same caliber, temporal proximity, and spatial neighborhood. Anomaly pruning truncates samples in the value fields that exceed the legal range or statistical threshold.

[0007] Furthermore, the quality control index matrix includes: The quality control index matrix includes evaluation unit identifier, index field and value field, and carries quality markers, which include missing markers, imputation source markers, pruning markers and caliber conversion markers.

[0008] Furthermore, the process of performing entropy and difference measurement calculations includes: The process includes: standardization of direction, mapping negative indicators to positive indicators, and scaling the value fields to zero to one, with the standardization based on linear scaling of the maximum and minimum values ​​within the same indicator field; entropy calculation, which calculates the entropy value of each indicator field based on the standardized value fields, including ratioization, i.e., proportionalization of the standardized value fields to obtain the ratio field; and difference measurement generation, which generates difference measurement fields based on the entropy values, constructed using a reverse discrete representation of entropy values ​​to minimize difference measurement when the entropy values ​​are close to a uniform distribution.

[0009] Furthermore, the process of generating the adjacent importance ratio includes: Stable term addition: Preset non-negative stable term values ​​are superimposed onto the difference measurement fields of adjacent preceding and following indicator fields before ratioization; Approach threshold backoff: When the difference between the difference measurement fields of adjacent preceding and following indicator fields is less than a preset approach threshold, the adjacent importance ratio field is assigned a value of one; Upper and lower limit truncation: The adjacent importance ratio field is restricted to between a preset lower limit and a preset upper limit; and monotonic constraint: The sequence of adjacent importance ratio fields is scanned and adjusted to ensure consistency with the sorting relationship of the difference measurement field.

[0010] Further, the G1 reverse recursive generation of weight vectors is performed. The process of generating weight vectors includes: The system employs hierarchical grouping, categorizing indicator fields into three groups: disaster-causing factors, disaster-inducing environment, and disaster-bearing bodies. Within each group, the system sorts the indicator fields in descending order of difference metrics. A G1 reverse recursive approach is used, recursively calculating weight values ​​from the last item forward along a deterministic order, utilizing the adjacent importance ratio field. Non-negative normalization and consistency checks are performed, converting the weight value fields to non-negative values ​​with a sum of one, and executing a monotonic check on the weights within each group. Finally, a weight vector is generated, containing both indicator fields and weight value fields.

[0011] Further, a weighted summation process is performed to generate a comprehensive risk score sequence. The process of generating the comprehensive risk score sequence includes: The process includes: preparing numerical data with the same caliber; scaling and mapping the value fields in the quality control indicator matrix based on the maximum and minimum value records in the processing log to generate weighted value fields; performing weighted summation on each indicator field within the same evaluation unit identifier and summing them up; and boundary verification to determine whether the comprehensive risk score value field falls within the preset scoring range; and generating a comprehensive risk score sequence, which includes the evaluation unit identifier field and the comprehensive risk score value field.

[0012] Furthermore, the process of generating five-level risk labels and parameter records by performing K-means clustering includes: K-means clustering is initialized multiple times. Under the premise of a fixed number of five clusters, the clustering iteration is repeatedly performed using different initialization numbers and random number seeds. Silhouette coefficient evaluation is performed by calculating the silhouette coefficient based on the distance between samples within and between clusters, and the clustering result with the largest silhouette coefficient is selected as the optimal result. Five-level risk labels are generated by sorting the cluster centers of the optimal result according to their center values ​​and mapping them to five discrete levels. Five-level risk labels and parameter records are generated. The five-level risk labels include an evaluation unit identifier field and a five-level risk label field. The parameter records include the initialization number, random number seed, cluster center, and silhouette coefficient.

[0013] The key innovations of this invention include: (1) In the entropy difference set obtained from the quality control index matrix, the adjacent importance ratio is obtained from the ratio of adjacent difference measures, and the stable term is added during the generation of the adjacent importance ratio. At the same time, the monotonic constraint, the upper and lower limit truncation and the close threshold backoff are executed to form the adjacent ratio mark set with generation trajectory.

[0014] (2) The adjacent ratio mark set is grouped by level and sorted in descending order of the difference measure within the group to obtain a deterministic order. When the difference measure is the same within the same group, the deterministic order is determined by combining the missing proportion corresponding to the missing mark in ascending order, and the weight vector is generated by G1 reverse recursion according to the deterministic order.

[0015] (3) Perform non-negative normalization and consistency verification on the weight vector, wherein the consistency verification includes weight non-negative verification, weight sum one verification and intra-group weight monotonic verification, and associate the adjacent ratio label set, the weight vector and the comprehensive risk score sequence as the components of the result data package with the quality label and the parameter record in the GIS association visualization link.

[0016] The following are its main beneficial effects: (1) In view of the problem that the importance ratio of adjacent indicators in the existing method depends on subjective setting and is difficult to recalculate under the fluctuation of multi-source indicator data, the generation link of the adjacent importance ratio is organized in a regular way by the stability term, the monotonic constraint, the upper and lower limit truncation and the near threshold backoff, so that the adjacent importance ratio has a repeatable running path driven by the entropy difference set, and the generation trajectory is recorded by the adjacent ratio mark set, which facilitates the review and traceability under the same quality control indicator matrix input conditions.

[0017] (2) To address the problem that the ranking of existing methods is easily affected by implementation details when the contribution of indicators is close, making it difficult to reproduce the weight results, the weight generation process is limited to a clear group boundary by the hierarchical grouping and the deterministic order. Under the same difference measurement, the missing ratio corresponding to the missing label is introduced as the adjudication basis, so that the input order of the G1 reverse recursion has a consistent deterministic rule, thereby forming a cross-step comparable link between the weight vector generation process and the adjacent ratio label set and the quality label.

[0018] (3) In view of the problem that the weight output lacks verifiable constraints in the existing method and is prone to inconsistency in the subsequent comprehensive risk scoring and classification, the availability conditions of the weight vector are made explicit through the non-negative normalization and consistency verification, and the weight vector, the comprehensive risk score sequence, the quality label and the parameter record are included in the result data package, so that the data elements on which GIS association visualization depends have a unified recording standard and traceability relationship, thereby supporting the recalculation and verification of the five-level risk label and risk zoning map under the same input link. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating a multi-source data entropy correction G1 weighted flood risk zoning method provided in an embodiment of this application. Detailed Implementation

[0020] Example 1: Refer to Figure 1 This is a flowchart illustrating a multi-source data entropy correction G1 weighted flood risk zoning method provided in an embodiment of the present invention. The process may include at least steps S100-S800: S100: Obtain the zoning task identifier, evaluation unit version identifier and data source version identifier, perform data acquisition and metadata management processing, and generate the original indicator set; S200: Based on the original indicator set, perform quality control and caliber alignment processing to generate a quality control indicator matrix; S300: Based on the quality control index matrix, perform entropy value and difference measurement calculation and generate entropy difference set; S400: Based on the entropy difference set, perform adjacent importance ratio generation processing to generate an adjacent ratio label set; S500: Based on the adjacent ratio label set and entropy difference set, perform G1 reverse recursive generation of weight vector processing to generate weight vector; S600: Based on the weight vector and quality control index matrix, perform weighted summation to generate a comprehensive risk score sequence; S700: Based on the comprehensive risk score sequence, perform K-means clustering to generate five-level risk labels, and generate five-level risk labels and parameter records; S800, based on five-level risk labels and parameter records, performs GIS-related visualization to generate result data packages, and generates result data packages and risk zoning maps.

[0021] S100: Obtain the zoning task identifier, evaluation unit version identifier and data source version identifier, perform data acquisition and metadata management processing, and generate the original indicator set; Specifically, in the business scenario of disaster risk assessment and geographic information data processing, the data acquisition and metadata management module is deployed to execute this step. This module receives the zoning task identifier, evaluation unit version identifier, and data source version identifier as trigger conditions. These trigger conditions are written into the execution queue by zoning task creation events, data source update events, or evaluation unit boundary update events and drive execution. The data acquisition and metadata management module organizes data objects for public service data processing tasks, ensuring that subsequent steps meet the requirements of a traceable data link. During execution, the module establishes a version number, source identifier, and collection time stamp for each type of input, and writes the version number into the version field of the original indicator set. This version field serves as the basis for caliber conversion and field verification in the subsequent S200 spatiotemporal caliber alignment processing.

[0022] Specifically, the evaluation unit boundary refers to the boundary expression and identification information of the evaluation unit in the spatial domain, including three parts: evaluation unit identifier, boundary geometry, and spatial reference information. The boundary geometry uses a set of surface features to represent the outer contour of the administrative region or watershed division. The spatial reference information records the coordinate datum, projection information, and boundary data version. After the data acquisition and metadata management module accesses the evaluation unit boundary data from the Geographic Information System (GIS) boundary database, it first reads the evaluation unit identifier and performs a uniqueness check. Records with duplicate, missing, or empty evaluation unit identifiers are written as abnormal records and set to a pending manual verification status. Then, a topological consistency check is performed on the boundary geometry, including surface feature closure checks, overlap / conflict checks, and hole anomaly checks. Detected geometric anomalies are written as boundary anomaly markers, and the original geometric snapshot is retained. Finally, the spatial reference information is read for consistency and written to the spatial reference field, which includes the coordinate datum name, projection name, and boundary data version number. The evaluation unit boundary is formed into an evaluation unit boundary data packet in this step. The evaluation unit boundary data packet is loaded as a boundary field in the original index set and serves as the input source for spatial boundary clipping or regional statistics in subsequent S200.

[0023] Specifically, the multi-source indicator data refers to a collection of tabular indicator data, raster indicator data, and time-series indicator data. Tabular indicator data comes from statistical yearbooks, bulletins, or business ledgers and includes evaluation unit association fields, indicator fields, value fields, and time stamp fields. Raster indicator data comes from remote sensing or grid products and includes raster coverage, resolution, pixel value fields, and time stamp fields. Time-series indicator data comes from monitoring stations or business monitoring platforms and includes station identifier, observation value fields, timestamp fields, and station spatial location fields. The data acquisition and metadata management module performs access standardization processing on multi-source indicator data. This standardization process includes data source connection, file parsing, field extraction, and structured encapsulation: When accessing tabular indicator data, it reads indicator fields and performs field dictionary mapping, mapping data source field names to unified indicator field names. It also reads value fields and performs data type validation, writing non-numeric, empty strings, and illegal characters into value anomaly markers while retaining the original text. When accessing raster indicator data, it reads raster coverage and resolution and writes them into the spatial aperture field. It encapsulates the pixel value field into a raster value carrier according to the original storage order and writes missing carrier markers for missing pixels. When accessing time-series indicator data, it reads the timestamp field and performs timestamp validity checks, writing invalid timestamps into time anomaly markers. Simultaneously, it reads the site spatial location field and writes it into the spatial positioning field. The access normalization process in this step does not perform caliber conversion, spatial pruning, interpolation, or clipping. It only completes the original access, field extraction, and anomaly recording of multi-source indicator data, so that the multi-source indicator data retains the original value semantics and enters the indicator data field of the original indicator set. The indicator data field serves as the input source for spatiotemporal caliber alignment, missing data gating, hierarchical interpolation, and anomaly clipping in subsequent step S200.

[0024] Furthermore, the spatiotemporal metadata refers to a set of metadata describing the statistical period, time granularity, spatial resolution, statistical scope, and unit scope of the multi-source indicator data. It includes two parts: time-based metadata and spatial-based metadata. The time-based metadata includes the statistical period type, start and end time, time granularity identifier, and data release period identifier. The spatial-based metadata includes the spatial resolution identifier, spatial coverage identifier, administrative division version identifier, and unit scope identifier. The data acquisition and metadata management module extracts spatiotemporal caliber metadata from the data source header information, product description fields, and caliber dictionary, and performs a legality check on the unit caliber identifier. The legality check includes unit name verification and caliber description completeness verification. Records with missing unit names or missing caliber descriptions are marked with a caliber missing tag while retaining the original description text. At the same time, the data acquisition and metadata management module establishes an association index between the spatiotemporal caliber metadata and the multi-source indicator data. The association index includes the data source identifier, indicator fields, and spatiotemporal caliber metadata identifier. The association index is written into the caliber field of the original indicator set and used in subsequent steps S200 for time granularity unification, unit caliber conversion, and field verification.

[0025] In the engineering implementation, the evaluation unit boundary adopts a set of county-level administrative area features. The multi-source indicator data includes rainfall raster products, topographic raster products, population and asset statistical tables, and time-series water level data from monitoring stations. The spatiotemporal metadata includes the annual scope of the statistical yearbook, the spatial resolution scope of the raster products, and the minute-level time scope of the monitoring stations. Upon access, the data acquisition and metadata management module writes the evaluation unit identifier of the county-level features into the boundary field, the coverage and resolution of the rainfall raster into the spatial scope field, the statistical period and unit name of the yearbook tables into the time-scope metadata and unit scope identifier, and writes the timestamp validity check results of the monitoring station time series into the time anomaly flag, thereby obtaining the structured original indicator set.

[0026] Finally, the data acquisition and metadata management module merges and encapsulates the evaluation unit boundary, multi-source indicator data, and spatiotemporal caliber metadata to generate the original indicator set. The original indicator set includes an evaluation unit identifier field, a boundary geometry field, an indicator field, a value field, a time stamp field, a spatial location field, a data source identifier field, a version field, and a caliber field. The evaluation unit identifier field and the boundary geometry field are used for spatial boundary clipping or regional statistics in subsequent S200 inputs. The indicator field and the value field are used for missing data gating, hierarchical interpolation, and anomaly clipping in subsequent S200 inputs. The time stamp field, the spatial location field, and the caliber field are used for spatiotemporal caliber alignment and unit caliber conversion in subsequent S200 inputs.

[0027] In summary, this step achieves the following technical results: It integrates evaluation unit boundaries, multi-source indicator data, and spatiotemporal metadata into a unified raw indicator set, incorporating source identifiers, version fields, and caliber fields into the same data object. This raw indicator set undergoes structured encapsulation and anomaly logging, forming the input basis for subsequent S200 operation. The raw indicator set maintains consistency in field names and version association indexes across steps, supporting traceability processing in subsequent links.

[0028] S200: Based on the original indicator set, perform quality control and caliber alignment processing to generate a quality control indicator matrix; Specifically, after the data acquisition and metadata management module completes S100, the quality control and caliber alignment module receives the original indicator set as input for this step and starts running. The start is triggered by the zoning task identifier, which carries the target time granularity identifier and the target spatial caliber identifier and writes them into the running queue. The quality control and caliber alignment module reads the evaluation unit identifier field, boundary geometry field, indicator field, value field, time stamp field, spatial positioning field, data source identifier field, version field, and caliber field from the original indicator set, and establishes a processing serial number for each indicator record in the same processing session. The processing serial number is used to record the execution status of caliber alignment, missing gating, hierarchical imputation, and abnormal pruning within this step. The spatiotemporal caliber alignment refers to the consistency processing of the time stamp field, spatial positioning field, and caliber field in the original indicator set under the constraints of the target time granularity identifier and the target spatial caliber identifier. The missing gating refers to the determination of the missing status and missing ratio of the value field and the triggering of subsequent imputation or deletion actions. The hierarchical imputation refers to the generation of replacement values ​​for missing items and the recording of the source level according to the priority of same source and same caliber, temporal proximity, and spatial neighborhood. The abnormal pruning refers to the truncation processing of samples in the value field that exceed the legal range or statistical threshold and the recording of pruning information. The quality mark refers to the mark set corresponding to each value field. The quality mark includes missing mark, imputation source mark, pruning mark, and caliber conversion mark and is written into the output of this step, so that when S300 uses the quality control indicator matrix as input, it has a traceable running trace.

[0029] Specifically, the spatiotemporal alignment process is completed within the same evaluation unit identifier range. The quality control and alignment module first extracts temporal and spatial metadata based on the caliber field, and binds its data source identifier field and version field to each indicator field to form a caliber alignment configuration. Then, it performs unified time granularity processing, mapping the time marker field to the target time granularity identifier. The mapping follows the statistical period type in the time caliber metadata. For tabular indicator records, period boundary alignment is used while keeping the value field unchanged. For time-series indicator records, window aggregation is performed according to the target time granularity identifier. The window aggregation uses window accumulation when the caliber field is labeled as cumulative, window mean when the caliber field is labeled as mean, and window maximum or minimum value when the caliber field is labeled as extreme value. Records that cannot meet the window aggregation conditions are marked with a time anomaly and the original value field is retained. The spatial caliber alignment process is completed based on the boundary geometry field and the spatial positioning field. For raster-type index records, the quality control and caliber alignment module spatially overlays the raster coverage area with the boundary geometry field, performs in-boundary clipping, and performs region statistics on the clipped pixels. The region statistics generate area statistics when the caliber field is labeled as an area type, and generate mean statistics when the caliber field is labeled as an intensity type, and write the generated statistics back to the value field. For station-type time series records, the quality control and caliber alignment module maps the spatial positioning field to the inclusion relationship of the boundary geometry field, determines the corresponding evaluation unit identifier, and writes the mapping relationship into the spatial mapping record. If the station spatial positioning field falls into multiple boundary geometry fields or does not fall into any boundary geometry field, a spatial anomaly marker is written and the record is set to a pending verification state. The unit caliber conversion is performed when the caliber field provides a unit caliber identifier. The quality control and caliber alignment module reads the conversion relationship from the unit dictionary, performs numerical conversion on the value field according to the conversion relationship, and writes the conversion mark. The conversion mark records the unit name before conversion, the unit name after conversion, and the conversion version number. For records where the unit caliber identifier is missing or the unit dictionary lacks a matching item, a missing caliber mark is written while the original value field is preserved. The missing caliber mark is output along with the quality mark for reference in subsequent steps. The above processing forms a consistent data representation in the public service data processing task chain, ensuring that the original indicator set has a unified time stamp field and evaluation unit identifier attribution relationship before entering the subsequent missing gate and anomaly pruning.

[0030] Furthermore, the missing data gating and hierarchical imputation are performed on the evaluation unit identifier dimension according to the indicator fields. The quality control and caliber alignment module calculates the missing ratio for each indicator field. The missing ratio is determined by the number of missing records and the total number of records for that indicator field across all evaluation unit identifiers. The number of missing records is counted for records whose value fields are null, invalid, or invalidated by time anomaly or spatial anomaly markers. When the missing ratio is greater than a preset threshold, the indicator field is marked as unavailable and the column corresponding to the indicator field is deleted when generating the quality control indicator matrix. At the same time, the name of the unavailable indicator field and the missing ratio are recorded in the processing log. For indicator fields with a missing proportion less than a preset threshold, the quality control and caliber alignment module writes a missing marker to the missing item and triggers hierarchical imputation. The minimum set of hierarchical imputation includes three layers of rules: priority for same source and same caliber, temporal proximity order, and spatial neighborhood order. In the priority layer for same source and same caliber, the quality control and caliber alignment module searches for adjacent time marker field records of the same evaluation unit identifier under the constraints of the same data source identifier field and the same caliber field, selects a substitute value according to the principle of minimum time distance, and writes it into the value field, while writing the imputation source marker as same source and same caliber. In the temporal proximity order layer, when the priority layer for same source and same caliber does not find a suitable replacement value, the module selects a substitute value according to the principle of minimum time distance and writes it into the value field, while writing the imputation source marker as same source and same caliber. When recording, the quality control and caliber alignment module performs interpolation based on the preceding and succeeding neighboring records of the time stamp field under the same evaluation unit identifier. The interpolation uses the mean of neighboring records or the value of the preceding record, and the interpolation source marker is written as time-series neighbor. At the spatial neighborhood order layer, if the time-series neighbor order layer still cannot generate a substitute value, the quality control and caliber alignment module generates spatial adjacency relationships based on the boundary geometry field. The spatial adjacency relationships are determined by the boundary geometry shared boundary, and the available values ​​of the same indicator field of adjacent evaluation unit identifiers are retrieved in the spatial adjacency relationships. The neighborhood median is written into the value field, and the interpolation source marker is written as spatial neighborhood. For missing items that still cannot be interpolated, the missing marker is kept valid and written into the interpolation failure record. The interpolation failure record is associated with the processing serial number, so that subsequent steps can identify the source status of the value field when reading the quality marker.

[0031] Furthermore, the anomaly pruning is performed after spatiotemporal alignment and hierarchical interpolation are completed. The anomaly pruning is performed on the indicator field dimension, and the value field of each evaluation unit identifier is judged one by one. The quality control and caliber alignment module first screens outliers based on the legal range of the indicator field. The legal range is given by the business boundary description and field verification rules in the caliber field. Records exceeding the legal range are marked with a pruning flag and enter truncation processing. Then, the anomaly threshold is calculated based on the median and interquartile range. The anomaly threshold is calculated within the range of all evaluation unit identifiers in the same indicator field. The interquartile range is determined by the difference between the upper quartile and the lower quartile. The anomaly threshold is determined by the interquartile range of the lower quartile minus a preset multiple and the interquartile range of the upper quartile plus a preset multiple. For samples exceeding the anomaly threshold, truncation processing is performed. The truncation processing replaces the value field with the corresponding threshold boundary value and writes the pruning flag into the truncation type, threshold version number, and processing serial number. For records with both missing and trimmed markers, the quality control and caliber alignment module retains both the interpolation source marker and the trimmed marker, allowing the source link of the same value field to be traced by subsequent steps. For records with caliber conversion markers, the trimming decision is executed based on the converted value field, and the caliber conversion marker version number is recorded in the trimmed marker as an associated field.

[0032] In the engineering implementation, the boundary of the evaluation unit is a county-level administrative region, the original indicator set includes disaster-bearing body indicators from the annual statistical table, disaster-causing factor indicators from the monthly rainfall grid, and disaster-inducing environment indicators from the station's minute-level water level time series. The quality control and caliber alignment module receives the target time granularity identifier as an annual scale and the target spatial caliber identifier as a county-level scale. The spatiotemporal alignment process aggregates monthly rainfall rasters into annual scale windows and performs regional statistics within the boundary geometry field to generate annual rainfall statistics. It also aggregates station water level time series into annual water level statistics and maps them to county-level evaluation unit identifiers. Subsequently, it calculates the missing proportion according to the indicator fields and triggers hierarchical imputation for fields with low missing proportions. Imputation is completed by selecting adjacent year records in the same source and same caliber priority layer, selecting previous year records in the temporally adjacent order layer, and selecting the median of adjacent counties in the spatial neighborhood order layer. The imputation source marker is written to the same source and same caliber, temporally adjacent, or spatial neighborhood, respectively. Then, it performs legal range screening and abnormal threshold truncation based on median and interquartile range on the imputed value fields and writes the truncation marker into the truncation information and threshold version number to form an input data structure that can be directly entered into S300.

[0033] Finally, the quality control and caliber alignment module aggregates and generates the quality control index matrix within the same processing session. The quality control index matrix includes evaluation unit identifiers, index fields, and value fields, and establishes a one-to-one correspondence between the quality markers and the value fields according to the processing serial number. The quality markers include missing markers, imputation source markers, pruning markers, and caliber conversion markers, and are output along with the quality control index matrix. Among them, the evaluation unit identifier serves as the row index input for the unified and standardized execution direction of subsequent S300, the index field serves as the column index input for subsequent S300 to calculate entropy values ​​and difference measures, the value field serves as the numerical input for subsequent S300 to calculate entropy values ​​and difference measures, and the quality marker serves as the auxiliary input for subsequent S300 to read data source status and identify imputation and pruning traces.

[0034] In summary, this step achieves the following technical effects: Based on the original indicator set, it aligns spatiotemporal scope and solidifies unit conversion information into scope conversion tags. Missing items are gating and hierarchical imputation are performed, and imputation source tags are written. Outliers are pruned, and pruning tags are written, ensuring that the quality control indicator matrix includes evaluation unit identifiers, indicator fields, value fields, and their corresponding quality tags. This quality control indicator matrix serves as input to S300, providing a traceable operational basis for subsequent entropy and difference measurement calculations.

[0035] S300: Based on the quality control index matrix, perform entropy value and difference measurement calculation and generate entropy difference set; Specifically, after completing S200, the entropy and difference measurement calculation stage receives the quality control index matrix as input for this step and enters the running state. This running state is triggered by the zoning task identifier, which carries a data version field and is bound to the processing serial number. The quality control index matrix already contains an evaluation unit identifier, an index field, and a value field, and carries the quality marker and the processing log along with the value field. This step first performs field verification on the quality control index matrix. The field verification checks whether there are duplicate or conflicting records in the value field according to the combination relationship between the evaluation unit identifier and the index field. Conflicting records are merged by first reading the record with the latest processing serial number, and the merging action is written to the processing log. For records with a valid missing marker and an empty value field, the missing marker is retained, and the record is set to a status marker that does not participate in the statistics. This status marker is written along with the processing log for subsequent entropy calculation stage reading. The process then proceeds to the directional unification and standardization link. Directional unification and standardization refers to the consistent expression of the positive meaning of different indicator fields and the completion of same-scale mapping. Directional unification refers to mapping negative indicators to positive indicators. Standardization refers to scaling the value fields of each indicator field to zero to one and writing the maximum and minimum values. The maximum and minimum values ​​are obtained by statistical analysis of the entire set of evaluation unit identifiers within the same indicator field and the same data version field, and are fixed as processing log fields. The processing log is continuously updated in this step and associated with the evaluation unit identifier and the indicator field.

[0036] Understandably, the directional standardization is performed column-by-column across indicator fields. The processing entity first reads the directional attribute of each indicator field from the processing log. This directional attribute is defined by the field conventions of the risk indicator system and remains consistent in the field verification records written into the processing log during the quality control phase. When the directional attribute indicates that the indicator field is a negative indicator, a maximum-minimum inverse mapping is performed. This maximum-minimum inverse mapping reads the maximum and minimum values ​​within the same indicator field and performs an inverse transformation on the value field identified by each evaluation unit, ensuring that the transformed value field satisfies the unified semantics that an increase in value corresponds to an increase in risk contribution. During the inverse transformation process, records with valid missing markers are not written with numerical values; only a directional standardization marker is written, and the missing status is retained in the processing log. For indicator fields indicating positive directional attributes, the original value field is retained, and a directional standardization marker is written. After direction unification is completed, standardization processing begins. This process reads the direction-unified value field within the same indicator field and performs a validity check. This check includes verifying whether the value field falls within the business boundary of the processing log record. The business boundary originates from the caliber field and the pruning mark record. For records with pruning marks, standardization directly reads the pruned value field while keeping the pruning mark unchanged. Subsequently, the maximum and minimum values ​​after direction unification are calculated within the same indicator field. These maximum and minimum values ​​are obtained after removing records with valid missing marks and empty value fields. These maximum and minimum values ​​are then written into the corresponding indicator field entry in the processing log. When the maximum and minimum values ​​are the same, resulting in a zero scale span, all standardized value fields of the indicator field are written to the same constant between zero and one, and a standardization degradation mark is added. This standardization degradation mark is written along with the processing log for subsequent entropy value statistics to identify the discrete state of the indicator field. For index fields with non-zero scale span, the value fields after unifying the direction are linearly scaled according to the maximum and minimum values ​​to obtain standardized value fields. The standardized value fields, together with the evaluation unit identifier and index fields, form a standardized matrix view, and are aligned with the quality mark according to the processing serial number.

[0037] After obtaining the standardized value field, this step proceeds to the calculation chain of entropy and difference measure. The entropy value is used to characterize the uniformity of distribution of the same indicator field across the evaluation unit identifier dimension, and the difference measure is used to characterize the discrimination magnitude corresponding to the entropy value and serves as the input source for the subsequent S400 generation of adjacent importance ratios. Specifically, for each indicator field, the processing entity extracts the standardized value field of the indicator field on each evaluation unit identifier from the standardized matrix view, and removes records marked as valid for missing data to ensure that the sample set participating in the statistics is consistent with the missing data gating result. Subsequently, the standardized value field participating in the statistics is subjected to proportioning processing. Proportioning processing obtains the proportion field by ratioing the standardized value field of each evaluation unit identifier to the sum of the standardized value fields of the indicator field, and records the source range and sample number of the sum value in the processing log. The entropy value of the indicator field is calculated based on the percentage value field. During the entropy calculation process, records with a percentage value of zero are marked with a zero percentage and logarithmic operations are skipped. This zero percentage mark is recorded in the processing log. The entropy value is output, written to the entropy value field, and bound to the indicator field. Subsequently, a difference measure field is generated based on the entropy value field. This difference measure field is constructed using a reverse discrete representation of the entropy value, ensuring that the difference measure is small when the entropy value is close to a uniform distribution and large when the entropy value deviates from a uniform distribution. The difference measure field is then written to the corresponding indicator field entry in the processing log. For indicator fields with standardized degradation marks, the entropy value field is written with a degradation entropy mark, and the difference measure field is set to zero. The degradation entropy mark and the standardized degradation mark are associated in the processing log, allowing subsequent steps to identify how the indicator field participates in sorting and adjacent ratio calculations.

[0038] In the engineering implementation, the evaluation unit identifier corresponds to the county-level administrative unit. The quality control index matrix has already completed the writing of caliber conversion markers, interpolation source markers, and pruning markers. In this step, after identifying a disaster-bearing body type index field as a negative index, the maximum-minimum reverse mapping is performed, the direction unification marker is written, and then a standardized value field is generated according to the maximum and minimum values ​​and written to the processing log. If a consistent disaster factor type index field is identified as a positive index, it directly enters the standardization process. Subsequently, the standardized value fields of the above index fields are subjected to proportion processing, and the entropy value field and difference measurement field are calculated. The entropy value field and difference measurement field, together with the evaluation unit identifier, index field, and standardized value field, are jointly encapsulated to generate an entropy difference set. The entropy difference set also carries the maximum value, minimum value, standardized degradation marker, and zero proportion marker recorded in the processing log, so that subsequent steps can maintain the traceability of the calculation path under the same data version field when reading the entropy difference set.

[0039] Finally, this step writes the generated entropy difference set as the output product into the result storage. The entropy difference set includes the evaluation unit identifier, index field, standardized value field, entropy value field, difference measurement field, and carries the processing log entries. The difference measurement field serves as the numerical input for S400 to generate the adjacent importance ratio, and the difference in the difference measurement field serves as the judgment input for S400 to perform near threshold backoff. The maximum and minimum values ​​in the processing log serve as the association input for S400 to trace the source path of the adjacent ratio, thereby completing the cross-step connection from S300 to S400.

[0040] In summary, this step unifies the direction of the quality control indicator matrix and solidifies standardized parameters, ensuring consistent semantics in the numerical expressions of each indicator field. Subsequently, entropy and difference measurement fields are formed and encapsulated into an entropy difference set, allowing the source link and data status of the difference measurement to be retained in the processing log. This entropy difference set is then passed to S400, providing a traceable data foundation for the generation of adjacent importance ratios.

[0041] S400: Based on the entropy difference set, perform adjacent importance ratio generation processing to generate an adjacent ratio label set; Specifically, after S300 is completed, this step is triggered into the running state by the zoning task identifier. The zoning task identifier carries a data version field and is bound to the processing serial number, allowing the calculation path for the same data version field in this step to be repeatedly invoked and tracked by the processing log. This step uses the entropy difference set as the input source. The entropy difference set includes at least an evaluation unit identifier, an indicator field, an entropy value field, and a difference measurement field, and is associated with the maximum and minimum values ​​in the processing log. This step first performs an input consistency check, which includes a completeness check of the indicator field set in the entropy difference set and a non-negativity check of the difference measurement field. For indicator fields with a degenerate entropy marker or a difference measurement field of zero, a degenerate participation marker is written and the record is retained. The degenerate participation marker is written with the processing log so that subsequent steps can identify the participation mode. Furthermore, this step constructs a carrier for generating adjacency relationships. These adjacency relationships do not directly depend on external sorting results, but are achieved by establishing associated records for two adjacent indicator fields in the adjacency ratio mark set. The generation of these associated records is based on the adjacency pair structure required for subsequent G1 reverse order recursion. In this step, candidate adjacency pairs are automatically generated according to the indicator field set under the same data version field. The generation rule for candidate adjacency pairs is to form an adjacency pair candidate sequence within the same indicator field set and write the adjacency pair generation record into the processing log. This allows subsequent steps to extract the corresponding adjacency pair records in that order when a deterministic order is obtained, without having to repeatedly construct the candidate structure.

[0042] After forming candidate adjacent pairs, this step enters the adjacent importance ratio generation chain. The adjacent importance ratio refers to the ratio result obtained by ratioing the difference measurement fields of two adjacent indicator fields, and establishes a binding relationship with the generation tag corresponding to the adjacent pair. Specifically, for each candidate adjacent pair record, this step reads the difference measurement fields of the adjacent preceding indicator field and the adjacent following indicator field respectively, and synchronously loads the missing tag statistical summary corresponding to the two indicator fields in the processing log during the reading stage, so that subsequent monotonic constraint processing has traceable reference information. Further, this step adds a stabilizing term before ratioing. The stabilizing term is a preset non-negative stabilizing term value. The stabilizing term value is one of the minimum sets of core parameters that suppress ratio fluctuations caused by excessively small difference measurement fields. The stabilizing term value is read from the zoning task configuration and written to the processing log by the parameter management unit when this step starts. The stabilizing term is added by superimposing the stabilizing term value onto the difference measurement fields of the adjacent preceding indicator field and the adjacent following indicator field respectively before performing ratioing. The ratioing output is written to the adjacent importance ratio field. Subsequently, a proximity threshold rollback is performed. This proximity threshold rollback is one of the minimum sets of core parameters, determined by the difference between the difference measurement fields of adjacent preceding and following indicator fields. When the difference is less than a preset proximity threshold, the adjacent importance ratio field is assigned a value of one, and a rollback flag is written into the generated flag. This rollback flag is associated with the adjacent generated records through a processing serial number. Further, upper and lower limit truncation is performed. This upper and lower limit truncation is one of the minimum sets of core parameters, implemented by restricting the adjacent importance ratio field to between a preset lower limit and a preset upper limit. For records where truncation occurs, a lower limit truncation flag or an upper limit truncation flag is written into the generated flag, and a summary of the value before truncation is written into the processing log for subsequent consistency checks. Subsequently, monotonic constraints are executed. These constraints apply sequence consistency constraints to adjacent importance ratio fields within the same deterministic order, ensuring that the ordering relationship between adjacent importance ratio fields and adjacent difference measurement fields remains consistent. In this step, monotonic constraints are implemented using candidate adjacent pair sequences as the basis, through pairwise scanning and write-back adjustments of adjacent importance ratio fields. When an adjacent importance ratio field is detected to have an inverse relationship with its preceding adjacent pair or to be less than one, the current adjacent importance ratio field is adjusted to the closest feasible value satisfying the monotonic relationship, and a monotonic adjustment flag is written into the generation flag. The monotonic adjustment flag and the upper and lower limit truncation flags can coexist, and their writing order is distinguished by the processing serial number. Understandably, the execution order of the aforementioned proximity threshold backoff, upper and lower limit truncation, and monotonic constraints is fixed as a step sequence field in the processing log, ensuring that the generation path of adjacent importance ratio fields for the same data version field remains consistent.

[0043] In the engineering embodiment, the entropy difference set comes from the flood risk zoning task at the county level, the evaluation unit identifier corresponds to the administrative division unit, the indicator field comes from multi-source indicator data and has been written with caliber conversion mark, interpolation source mark and pruning mark in S200, and the direction unification standardization is completed in S300 to generate the difference measurement field. In this step, after reading the difference measurement field of two adjacent indicator fields, the stable term value is added to generate the adjacent importance ratio field and written to the adjacent ratio mark set. At the same time, adjacent pairs with similar difference measurement fields are triggered to fall back to the threshold and write back mark; adjacent importance ratio fields that exceed the preset range after comparison value conversion are triggered to truncate the upper and lower limits and write truncation mark; adjacent importance ratio fields with reverse order in the same candidate adjacent pair sequence are triggered to monotonic adjustment mark and the adjusted adjacent importance ratio field is written back. Finally, this step outputs the adjacent ratio tag set, which includes at least the adjacent importance ratio field associated with two adjacent indicator fields and the generated tag. The adjacent ratio tag set is used as the input of the adjacent ratio tag set of S500, so that S500 can directly extract the corresponding adjacent pair records and perform G1 reverse recursion after obtaining the deterministic order. At the same time, this step records the versioned parameter summary of the stable item value, the preset lower limit value, the preset upper limit value and the preset close threshold in the processing log, so that the S800 result data package has an auditable parameter source link when aggregating the processing log.

[0044] In summary, the technical effects of this step are as follows: This step forms an adjacent importance ratio field based on the difference metric field of the entropy difference set, and suppresses ratio fluctuations caused by small difference metrics through the stable term value. The near-threshold backoff, upper and lower limit truncation, and monotonic constraints converge the abnormal forms of the adjacent importance ratio field to a traceable rule path, and leave verifiable traces through the generated tags. The adjacent ratio tag set is then connected to the deterministic order of S500 and the reverse recursion of G1, allowing the adjacent importance ratio field and generated tags output from this step to be directly reused in the weight vector generation stage.

[0045] S500: Based on the adjacent ratio label set and entropy difference set, perform G1 reverse recursive generation of weight vector processing to generate weight vector; Specifically, this step is triggered after the adjacent ratio tag set is output by S400. The triggering condition is that the adjacent ratio tag set and the entropy difference set are registered in the database under the same processing serial number, and the generation tags of the adjacent ratio tag set are not missing. The input sources of this step include the adjacent ratio tag set, the entropy difference set, and the processing log. The adjacent ratio tag set at least includes an indicator field, an adjacent importance ratio field, and generation tags. The entropy difference set at least includes an indicator field, an entropy value field, and a difference measurement field. The processing log at least includes a summary of the missing proportions and a summary of parameter records corresponding to the missing tags. This step first performs input association assembly, associating the adjacent ratio marker set with the entropy difference set according to the indicator field to generate a recursive record set. The recursive record set carries both the adjacent importance ratio field and the difference measurement field at the record level, and the association process is written to the association record field of the processing log. When it is found that there is an indicator field in the adjacent ratio marker set that is not registered in the entropy difference set, the indicator field is written to the association exception field of the processing log and a rollback recalculation is triggered. The rollback recalculation means keeping the processing serial number unchanged and restarting the S400 generation link. The trigger information of the rollback recalculation is also written to the processing log for subsequent audit and review.

[0046] Further, this step performs hierarchical grouping, which includes a disaster-causing factor group, a disaster-prone environment group, and a disaster-bearing body group, with each indicator field belonging to only one group. The implementation process of hierarchical grouping is as follows: the registration relationship between the indicator field and the group field is read from the processing log, and the registration relationship is bound to the recursive record set according to the indicator field to obtain a recursive candidate set with group identifiers; when the registration relationship is missing, the indicator field is temporarily stored as an ungrouped record and written to the missing group field of the processing log, and then the group identifier is filled in according to the registration order of the indicator field in the quality control indicator matrix and written to the filled record field of the processing log, so that the hierarchical grouping has a consistent landing path under the same data version field. After completing the hierarchical grouping, this step performs descending sorting of the difference metrics within each group. This descending sorting refers to arranging the indicator fields within the group from largest to smallest. When the difference metrics fields within the same group are the same, the order is determined by ascending sorting based on the missing percentage corresponding to the missing markers. The missing percentage is derived from the missing marker summary written in S200 and is queried from the processing log. When ties still exist, the indicator field registration order of the quality control indicator matrix is ​​used as the stable order and written into the parallel processing field of the processing log. Subsequently, this step concatenates the three group sorting results according to a preset group order to obtain a deterministic order. This deterministic order serves as the sole input order source for subsequent G1 reverse recursion, and a deterministic order record field is written into the processing log so that subsequent steps can directly reference it.

[0047] After obtaining the deterministic order, this step performs G1 reverse recursion to generate a weight vector. The G1 reverse recursion refers to the process of recursively calculating the weight values ​​item by item from the last item in the deterministic order. The adjacent importance ratio field required for the recursion comes from the adjacent ratio tag set and is obtained by retrieving adjacent indicator field pairs from the recursive record set. Specifically, this step first assigns a preset baseline weight value to the last indicator field of the deterministic order and writes this baseline weight value into the initial recursion value field of the processing log. Then, for each pair of adjacent indicator fields in the deterministic order, the adjacent importance ratio field corresponding to both is read, and the weight value of the preceding indicator field is recursively calculated according to the adjacent importance ratio field relative to the weight value of the following indicator field. During the recursion process, the generated tags are inherited synchronously, so that each weight value can be traced back to a specific record in the adjacent ratio tag set. When a pair of adjacent records has a rollback flag, truncation flag, or monotonic adjustment flag, this step writes the flag into the recursion basis field of the processing log and allows the adjacent importance ratio field to participate in the recursion. When a pair of adjacent records is missing or the generation flag indicator is unavailable, this step writes the adjacent pair into the recursion missing field of the processing log and triggers rollback recalculation. The rollback recalculation path is consistent with the aforementioned input association assembly stage. After completing the reverse recursion, an unnormalized weight sequence is obtained, and this unnormalized weight sequence is encapsulated into a weight vector according to the deterministic order. The weight vector contains at least an indicator field and a weight value field, and the encapsulation process is written into the weight encapsulation field of the processing log.

[0048] Subsequently, this step performs non-negative normalization and consistency verification on the weight vector. The non-negative normalization is a normalization process that converts the weight value fields into non-negative values ​​that sum to one. The process involves first performing a non-negative weight verification, writing records with negative weight values ​​to the non-negative weight verification field of the processing log, and then setting that weight value field to zero before participating in subsequent normalization. Next, a weight sum-to-one verification is performed, based on the sum of the weight value fields of the weight vector. If the sum does not meet the normalization condition, all weight value fields are scaled by the same ratio until the weight sum is one, and a summary of the scaling factor is written to the normalized record field of the processing log. Further, this step performs intra-group weight monotonic verification. This intra-group weight monotonic verification compares adjacent weight value fields pairwise in the deterministic order within the disaster-causing factor group, disaster-prone environment group, and disaster-bearing body group, determining whether the adjacent weights in the deterministic order satisfy the condition that the preceding weight is greater than the following weight. If this condition is not met, the corresponding adjacent indicator field is written to the monotonic verification anomaly field in the processing log, and a verification rollback is triggered. The verification rollback includes rereading the generation markers of the adjacent ratio marker set and prioritizing the selection of adjacent importance ratio fields that have not undergone truncation or monotonic adjustment, and re-evaluating them. The verification rollback process is written to the rollback record field in the processing log. If the verification rollback still does not meet the condition, the current weight vector is maintained, and the anomaly state is written to the processing log for reference and retention in subsequent steps. Understandably, the above non-negative normalization and consistency verification are all completed under the same processing serial number, and the parameter record summary under the same data version field is not rewritten; only the verification record written to this step is appended, forming a traceable versioned running trajectory.

[0049] In the engineering implementation, the evaluation unit identifier corresponds to the gridded evaluation unit within the county. The multi-source index data comes from hydrological station monitoring data, remote sensing raster data, and statistical summary data. The spatiotemporal caliber alignment, missing gating, hierarchical imputation, and anomaly pruning have formed the quality control index matrix in S200 and written missing markers, imputation source markers, pruning markers, and caliber conversion markers. The directional unification standardization and entropy value and difference measurement calculation have formed the entropy difference set in S300. The adjacent importance ratio generation has formed the adjacent ratio marker set in S400. This step reads the adjacent ratio marker set and groups the indicator fields into the disaster-causing factor group, disaster-prone environment group, and disaster-bearing body group according to the hierarchy. Within each group, the indicators are arranged in descending order according to the difference measurement field, and the parallel cases are handled to obtain the deterministic order. Then, G1 reverse recursion is performed according to the deterministic order to obtain the unnormalized weight sequence. The unnormalized weight sequence is encapsulated to form the weight vector, and non-negative normalization and consistency verification are performed. The weight vector is used as the weight vector input of S600, which is used by S600 to generate a comprehensive risk score sequence by weighted summation with the quality control indicator matrix. At the same time, the deterministic order record field, normalized record field, and monotonic verification anomaly field recorded in the processing log are converged in S800 to form the processing log component of the result data package.

[0050] In summary, this step achieves the following technical results: It establishes a deterministic order through hierarchical grouping and descending sorting by intra-group difference metrics, ensuring a verifiable sequence for the weight generation process. Based on the adjacent ratio marker set, a G1 reverse recursion is performed to generate a weight vector, with the recursion basis field and verification record field written to ensure traceability of the recursion path and parameter records. Non-negative normalization and consistency checks constrain the validity of the weight vector's numerical values ​​and the monotonic relationships within groups, and abnormal states are included in the processing log for reference in subsequent steps.

[0051] S600: Based on the weight vector and quality control index matrix, perform weighted summation to generate a comprehensive risk score sequence; Specifically, this step is triggered after the weight vector is output in S500, non-negative normalization is completed, and consistency verification is performed. The triggering condition is that the weight vector and the quality control index matrix are registered under the same processing serial number, and there are no unclosed loop records with weight non-negative verification anomalies in the processing log. The input sources for this step include the weight vector, the quality control index matrix, and the processing log. The weight vector contains at least an index field and a weight value field, and retains the index field arrangement relationship corresponding to the deterministic order in S500. The quality control index matrix contains at least an evaluation unit identifier, an index field, and a value field, and the quality mark is written in S200. The quality mark includes a missing mark, an imputation source mark, a pruning mark, and a caliber conversion mark. The processing log contains at least the maximum and minimum value records written in S300 and the caliber conversion mark records written in S200. This step first performs input assembly and consistency verification. Specifically, the weight vector is indexed by indicator fields, and the quality control indicator matrix is ​​aggregated into an evaluation unit record set by evaluation unit identifier. Then, the indicator fields in the evaluation unit record set are read one by one, and the weight value fields of the same indicator fields in the weight vector are retrieved and paired to form a weighted calculation record set. This weighted calculation record set carries the evaluation unit identifier, indicator field, value field, weight value field, and quality mark at the record level. When an indicator field in the quality control indicator matrix is ​​not registered in the weight vector, the indicator field and its corresponding evaluation unit identifier are written into the missing weight record field of the processing log, and the indicator field is removed from the weighted calculation record set according to the preset missing weight handling rules. When an indicator field in the weight vector is not present in the quality control indicator matrix, the indicator field is written into the missing value record field of the processing log, and the original weight vector record is retained but not participated in this weighted calculation, thus ensuring that the execution path of this step remains aligned with the closed loop of the S500 weight generation logic.

[0052] Further, this step performs a caliber-based numerical preparation for weighted summation on the value fields. This caliber-based numerical preparation refers to the process of unifying the value fields in the quality control index matrix that can be summed to a comparable scale. The processing basis comes from the maximum and minimum value records written by S300 and the caliber conversion flag records written by S200 in the processing log. Specifically, for each index field, its corresponding maximum and minimum value records in the processing log are read. A scaling mapping is performed on the value field to generate a weighted value field. When the caliber conversion flag in the quality flag indicates that the value field has undergone spatiotemporal caliber conversion, the converted value field is preferentially used in the scaling mapping. When the pruning flag in the quality flag indicates that the value field has undergone abnormal pruning, the pruned value field is still used in the scaling mapping, and a pruning participation record field is written in the processing log. When the missing flag in the quality flag indicates that the value field was obtained by hierarchical interpolation, the interpolated value field is still used in the scaling mapping, and the interpolation source flag is written in the interpolation participation record field of the processing log. Understandably, the above-mentioned numerical preparation with the same caliber is one of the core parameter minimal sets in this step. Its input depends only on the value field of the quality control index matrix and the maximum and minimum value records of the processing log, and the weight vector is not modified. As an optional extended function, this step can also encapsulate the quality mark along with the weight value field into a score traceability record for direct reference when the subsequent result data package is aggregated. However, this encapsulation does not change the generation rules of the comprehensive risk score sequence.

[0053] After generating the weighted value field, this step performs a weighted summation operation and generates a comprehensive risk score sequence. Specifically, for the same evaluation unit identifier, the weighted calculation record set corresponding to the evaluation unit is traversed by indicator field. For each record, the weighted value field and the weight value field are weighted and superimposed item by item, and the summation is performed within the evaluation unit to obtain the comprehensive risk score field of the evaluation unit. When there is an indicator field in the evaluation unit that has been removed by the missing weight handling rule, this step writes the removal information into the removal summary field of the processing log, and rescales the remaining weight value fields participating in the summation by the same proportion before participating in the summation, so that the total weight of the summation input between different evaluation units remains consistent. Furthermore, this step performs boundary verification on the comprehensive risk score field. Boundary verification refers to the process of determining whether the score field falls within a preset score range. When the score field exceeds the preset score range, the corresponding evaluation unit identifier is written into the score anomaly field of the processing log, and a backtracking check is triggered. The backtracking check includes rereading the value field of the evaluation unit in the quality control index matrix and the quality mark, and recalculating the weighted value field by comparing the caliber conversion mark record and the maximum and minimum value records in the processing log. If a score anomaly still exists after the backtracking check is completed, the score anomaly status is retained and kept as part of the processing log for reference in subsequent steps. Therefore, this step outputs the comprehensive risk score sequence, which includes an evaluation unit identifier field and a comprehensive risk score value field. The comprehensive risk score sequence is registered as the comprehensive risk score sequence input in S700 for subsequent K-means clustering operations. At the same time, the missing weight record field, missing value record field, imputation participation record field, pruning participation record field, and scoring anomaly field newly added in the processing log will be aggregated into the processing log component of the result data package in S800, and will form a cross-step connection with the subsequent GIS association visualization steps.

[0054] In the engineering embodiment, the boundary of the evaluation unit is the boundary of the gridded evaluation unit within the watershed. The multi-source index data includes station rainfall and water level observation data, remote sensing land cover data, and statistical summary data. The spatiotemporal caliber metadata records the temporal granularity, spatial resolution, and statistical caliber of each data source. After the quality control index matrix is ​​formed by S200 and written to the quality tag, and the weight vector is formed by S500, this step is automatically triggered by the risk zoning calculation unit. The quality control index matrix is ​​aggregated according to the evaluation unit identifier and the weight vector is associated according to the index field. The maximum and minimum value records in the processing log are read, and the value field is scaled and mapped to generate a weighted value field. Then, each index field in the same evaluation unit is weighted and superimposed item by item and summed to form a comprehensive risk score value field. Finally, the comprehensive risk score sequence is output and written to the version registration, so that the comprehensive risk score sequence can be directly called by S700 and retains the traceability relationship with the quality tag.

[0055] In summary, this step establishes a consistent pairing link between the weight vector and the quality control index matrix based on index fields, and records missing weights and values ​​in the processing log to create traceable records. By reading the maximum and minimum value records from the processing log, the value fields are prepared with consistent numerical standards, ensuring that the input for weighted summation meets the comparability requirements and remains consistent with the processing standards of previous steps. The output comprehensive risk score sequence is bound to evaluation unit identifiers at the field level and registered as input for subsequent clustering steps, while retaining scoring anomalies and participation records for use in results aggregation and audit review.

[0056] S700: Based on the comprehensive risk score sequence, perform K-means clustering to generate five-level risk labels, and generate five-level risk labels and parameter records; Specifically, this step is triggered after the S600 outputs the comprehensive risk scoring sequence and completes version registration. The triggering condition is that there are no unhandled status records corresponding to the scoring anomaly field in the processing log, and the comprehensive risk scoring sequence completes the evaluation unit identifier consistency check with the evaluation unit boundary under the same processing serial number. The input sources for this step include the comprehensive risk scoring sequence and the processing log, wherein the comprehensive risk scoring sequence at least includes an evaluation unit identifier field and a comprehensive risk scoring value field, and the processing log at least includes a serial number field and a data version field used to trace the generation caliber of the comprehensive risk scoring value field. To facilitate the spatial attachment of risk labels in subsequent GIS visualization, this step first performs clustering input preparation on the comprehensive risk score sequence. Clustering input preparation refers to the process of organizing the comprehensive risk score value fields into a score sample set that can be used for clustering operations. Specifically, this includes sorting and deduplicating by evaluation unit identifier field, removing records with empty score value fields and writing their evaluation unit identifier fields into the processing log (clustering exclusion field), and performing range verification on the score value fields and writing out-of-bounds records into the processing log (clustering anomaly field). When a record exists in the clustering anomaly field, this step isolates the score sample set corresponding to the evaluation unit identifier field according to preset anomaly handling rules, preventing it from participating in this round of clustering operations but still retaining it in the label completion link of the output stage. The label completion link will be explained later.

[0057] Furthermore, after constructing the scoring sample set, this step performs multiple initializations of K-means clustering. "Multiple initializations" refers to repeatedly executing the clustering iteration process using different initialization numbers and random number seeds for the cluster centers, while maintaining a fixed five-level cluster size. "Five levels" means that the final output risk labels are divided into five discrete levels, each corresponding one-to-one with the evaluation unit identifier field. In this step, "K-means clustering" is executed by the clustering calculation unit, which consists of an initialization module, a distance calculation module, a center update module, a convergence determination module, and a result evaluation module. Each module operates under controlled access according to a unified processing sequence number, and the parameters are written into the record after each run. Specifically, the initialization module extracts five initial cluster centers from the scoring sample set according to the initialization number and random number seed to form an initial center set. The distance calculation module calculates the distance between each scoring sample and the initial center set and generates a sample affiliation set. The center update module calculates the mean of the sample affiliation set for each cluster and updates the cluster centers to obtain an updated center set. The convergence determination module compares the change amplitude of two adjacent updated center sets and outputs the clustering result of this initialization when the change amplitude meets a preset convergence threshold or the number of iterations reaches a preset upper limit. The preset convergence threshold and the upper limit of the number of iterations constitute one of the minimum sets of core parameters used to maintain automated operation in this step. The minimum set of core parameters includes at least the number of clusters, the upper limit of the number of iterations, the preset convergence threshold, and the number of initializations. As an optional extension function, this step can use batch execution parallel scheduling for different initialization numbers and write the batch identifier of the parallel scheduling into the parameter record, but this extension does not change the evaluation criteria and selection rules of the clustering results.

[0058] Further, this step performs evaluation criteria calculation on the clustering results obtained from each initialization and selects the optimal result. The "evaluation criterion" is the silhouette coefficient, which is a numerical indicator characterizing cluster compactness and separation based on the intra- and inter-cluster distances of samples. In this step, the result evaluation module calculates the silhouette coefficient for the scored sample set and its corresponding sample affiliation set without introducing additional input fields and outputs the silhouette coefficient value field. Specifically, the result evaluation module calculates the average distance of each scored sample within its own cluster and generates an intra-cluster distance field, while simultaneously calculating its average distance to other clusters and generating an inter-cluster distance field. The sample silhouette value field is obtained from the intra- and inter-cluster distance fields, and the mean of the sample silhouette value fields of all scored samples is then calculated as the silhouette coefficient value field for this initialization. After each initialization, this step writes the initialization number, random number seed, cluster centers, and silhouette coefficient into the parameter record, and writes the order of the cluster centers and the sorting information of the scored samples into the clustering traceability field of the processing log. After all initializations are completed, this step selects the clustering result with the largest silhouette coefficient value from the candidate result set as the optimal result, and writes the initialization number corresponding to the optimal result into the optimal initialization field of the parameter record to form a recalcible result selection basis; when there are candidates with the same silhouette coefficient value, this step compares the separation of the cluster centers according to the preset parallel handling rules and selects the result with the larger separation, and writes the parallel handling process into the parallel handling field of the processing log.

[0059] After obtaining the optimal result, this step generates five-level risk labels and forms an output structure corresponding to the evaluation unit identifier field. The "five-level risk label" refers to the label field obtained by orderly mapping the five clusters in the optimal result according to their risk level. This step sorts the cluster centers of the optimal result by their center values ​​from smallest to largest to obtain a center sequence, and maps the position of the center sequence to the risk label position, thus enabling each scored sample to obtain a corresponding risk label field based on its cluster. The risk label field and the evaluation unit identifier field combine to form a five-level risk label output record set. Further, to directly interface with the S800 GIS-related visualization, this step sorts the five-level risk label output record set by the evaluation unit identifier field and removes duplicates, and writes the center sequence, center position mapping rules, and cluster centers into the cluster center field of the parameter record. For the evaluation unit identifier field that was isolated during the clustering input preparation stage, this step executes a label completion link. Specifically, it reads the comprehensive risk score value field and performs nearest center matching based on the center sequence to obtain the risk label field of the evaluation unit. This completion action is then written to the completion record field of the processing log. When the comprehensive risk score value field is empty and cannot be completed, the evaluation unit identifier field is written to the uncompleted field of the processing log. During output, the risk label field is set to empty, but the evaluation unit identifier field is retained and not deleted, so that the evaluation unit can still be located in subsequent result aggregation.

[0060] Finally, this step outputs the five-level risk labels and the parameter records. The five-level risk labels include at least an evaluation unit identifier field and a five-level risk label field, and are registered as the five-level risk label input in S800, used for spatial association with the evaluation unit boundaries. The parameter records include at least an initialization number, a random number seed, cluster centers, and silhouette coefficients, and are registered as the parameter record input in S800, used to retain traceable information about the clustering process in the result data package. Understandably, the output of this step is consistent with the comprehensive risk score sequence in S600 at the evaluation unit identifier field level, allowing subsequent steps to simultaneously reference the comprehensive risk score sequence and the five-level risk labels to form the attribute field set of the risk zoning map, and to incorporate the parameter records and the processing log into the processing log component of the result data package.

[0061] In the engineering embodiment, the comprehensive risk score sequence originates from the comprehensive risk score value field of the watershed gridded evaluation unit. The clustering calculation unit is automatically triggered to run according to the processing sequence number each time a flood risk zoning task is started. The number of initializations is written into the processing log by the task configuration and remains unchanged during operation. The clustering calculation unit repeatedly performs multiple initializations of K-means clustering iterations on the score sample set, calculates the silhouette coefficients one by one and writes them into the parameter record, and finally selects the clustering result with the largest silhouette coefficient as the optimal result. Based on the cluster center, a five-level risk label field is obtained, and a five-level risk label output record set corresponding one-to-one with the evaluation unit identification field is output. At the same time, a parameter record containing the initialization number, random number seed, cluster center and silhouette coefficient is output, so that S800 can directly use the five-level risk labels for spatial attachment and use the parameter record to reproduce the clustering process.

[0062] In summary, this step achieves the following technical results: By performing multiple initializations of K-means clustering on the comprehensive risk score sequence and using the silhouette coefficient as the evaluation criterion, the selection of clustering results is supported by recalculated parameter records. By generating five-level risk labels through ordered mapping of cluster centers, a stable correspondence is established between the risk labels and the evaluation unit identification field, facilitating subsequent GIS association. The output parameter records and processing logs form a traceability link for the clustering process, allowing the cluster centers, random number seeds, and optimal initialization numbers to be directly referenced during results aggregation.

[0063] S800, based on five-level risk labels and parameter records, performs GIS-related visualization to generate result data packages, and generates result data packages and risk zoning maps; Specifically, this step is triggered after the five-level risk labels are output in S700 and the parameter records are written. The triggering conditions include that the clustering operation status corresponding to the same processing serial number in the processing log is completed, and the evaluation unit identifier field in the five-level risk labels and the evaluation unit identifier field in the evaluation unit boundary have completed consistency verification. In addition to the five-level risk labels and the parameter records, the input sources of this step also include the quality label from S200, the weight vector from S500, and the comprehensive risk score sequence from S600. Simultaneously, the entropy difference set from S300 and the adjacent ratio label set from S400 are read, along with the caliber conversion label, the interpolation source label, the clipping label, and the processing log. The input data retains its original field names before entering the GIS association visualization. The GIS-related visualization is a geographic information system processing link, whose structure consists of a process of spatial data access, attribute field docking, symbolic rendering, layout output and result packaging. Spatial data access is used to load the boundary of the evaluation unit and generate a spatial index. Attribute field docking is used to write the five-level risk label, the comprehensive risk score sequence, the quality mark, the entropy difference set, the adjacent ratio mark set and the weight vector into the same evaluation unit attribute record according to the evaluation unit identifier field.

[0064] Furthermore, during the spatial data access phase, GIS-based visualization reads the boundary of the evaluation unit and parses its geometric shape field and evaluation unit identifier field. For duplicate evaluation unit identifier fields, a merging or retention strategy is executed. The merging or retention strategy is constrained by the boundary version field in the processing log and written into the boundary handling field of the processing log. When the evaluation unit boundary has an anomaly of empty geometry or self-intersection, GIS-based visualization writes the evaluation unit identifier field into the boundary anomaly field of the processing log and temporarily stores the evaluation unit boundary in a state to be repaired without blocking the subsequent attribute field docking process. Subsequently, the attribute field docking stage begins. GIS association visualization performs primary key docking based on the evaluation unit identifier field, writing the level 5 risk label field of the level 5 risk label into the evaluation unit attribute record. At the same time, it writes the initialization number, random number seed, cluster center, and silhouette coefficient related to the optimal result into the parameter record, and associates the parameter record with the processing serial number to form a traceable operation record. When the level 5 risk label has a missing evaluation unit identifier field, GIS association visualization determines the supplementation status based on the supplementation record field in the processing log. For evaluation unit attribute records that have not been supplemented and have empty labels, it retains the null value and writes it into the label missing field of the processing log, so that the evaluation unit identifier field can still be completely output in the subsequent result packaging stage.

[0065] Furthermore, in the symbolic rendering stage, GIS-related visualization reads the five-level risk label fields as the basis for classification rendering, and in the same evaluation unit attribute record, it reads the comprehensive risk score value field of the comprehensive risk score sequence as the annotation or grading verification field. The classification rendering rule is driven by the discrete values ​​of the five-level risk label fields. During the rendering process, the boundaries of evaluation units with the same label value are merged into the same risk layer, and the layer metadata is written into the rendering record field of the processing log. Furthermore, GIS-related visualization writes the missing marker, interpolation source marker, clipping marker, and caliber conversion marker in the quality markers into the quality information area of ​​the evaluation unit attribute record, and generates a visualization prompt field according to the quality information area. The visualization prompt field and the risk layer participate in map rendering together. When the quality marker indicates the existence of an abnormal handling record, GIS-related visualization writes the corresponding evaluation unit identifier field into the rendering record field, so that the risk zoning map retains quality clues when displayed. The process then proceeds to the layout output stage. GIS-based visualization overlays the risk layer with the base map elements to generate a risk zoning map. This risk zoning map includes spatial geometry and classification display information corresponding to the five-level risk labels. The processing serial number, boundary version field, and rendering record field are written into the map description information of the risk zoning map. When the boundary of an evaluation unit has a state that needs to be repaired, the risk zoning map is still output, but the boundary anomaly field count information is written into the map description information and simultaneously written into the processing log.

[0066] Finally, this step generates the output data package and completes the output field organization. The encapsulation process of the output data package consists of data aggregation, field verification, version registration, and storage writing. Specifically, in the data aggregation stage, the quality markers, entropy difference sets, adjacent ratio marker sets, weight vectors, comprehensive risk score sequences, five-level risk labels, and parameter records are summarized according to the evaluation unit identifier field. The caliber conversion markers, interpolation source markers, pruning markers, and processing logs are merged and written into the output metadata area to form a directly transferable output record set. In the field verification stage, field existence verification and null value placeholder verification are performed on the output record set. The verification rules are the field constraints recorded in the field specification of the processing log. When a missing field is found, the missing field name is written into the encapsulation exception field of the processing log, but the evaluation is not deleted. The evaluation unit identifier field; during the version registration stage, the processing serial number and boundary version field are written into the result metadata area and the version registration field of the processing log; during the storage writing stage, the risk zoning map is written into the map area of ​​the result data package, and the result record set is written into the data area of ​​the result data package, thereby outputting a result data package that simultaneously contains quality markers, weight vectors, and comprehensive risk score sequences, and includes evaluation unit identifiers, entropy difference sets, adjacent ratio marker sets, five-level risk labels, parameter records, caliber conversion markers, interpolation source markers, pruning markers, and processing logs in the same output. Understandably, the risk zoning map serves as a display carrier for terminal display and reading, and the result data package serves as a data carrier for storage medium writing and subsequent review process reading. The quality markers and processing logs provide field-level traceability entry points for the review process, and the parameter records provide clustering operation traceability entry points for the review process.

[0067] In the engineering implementation, the evaluation unit boundary adopts the township or grid evaluation unit boundary within the watershed. GIS visualization reads this boundary and connects it with the five-level risk label and the comprehensive risk score sequence according to the evaluation unit identification field, generating a risk zoning map covering the entire watershed and simultaneously generating a result data package. The result data package, under the same processing serial number, writes quality tags, weight vectors, and the comprehensive risk score sequence, while retaining the entropy difference set, the adjacent ratio tag set, and the parameter records. This data is used to display the risk zoning map on the emergency management business terminal and to write it into the archiving medium for business data storage. The processing log is output along with the result data package to locate boundary anomalies, missing tags, and encapsulation anomalies.

[0068] Summary of the technical effects of this step: This step uses GIS visualization to connect the five-level risk labels with the boundaries of evaluation units and generate a risk zoning map, establishing a stable correspondence between the zoning output and the evaluation unit identification fields. Through the result data package encapsulation process, quality markers, weight vectors, comprehensive risk score sequences, entropy difference sets, adjacent ratio marker sets, parameter records, and processing logs are aggregated into a single package, giving the output a unified data format. The version registration field, along with the processing serial number, connects the risk zoning map and the result data package, establishing a traceable link between the output content and the operational process.

Claims

1. A multi-source data entropy-corrected G1-weighted flood risk zoning method, characterized in that, include: S100: Obtain the zoning task identifier, evaluation unit version identifier and data source version identifier, perform data acquisition and metadata management processing, and generate the original indicator set; S200: Based on the original indicator set, perform quality control and caliber alignment processing to generate a quality control indicator matrix; S300. Based on the quality control index matrix, perform entropy value and difference measurement calculation processing to generate an entropy difference set; the entropy difference set includes evaluation unit identifier, index field, standardized value field, entropy value field and difference measurement field. S400. Based on the entropy difference set, perform adjacent importance ratio generation processing to generate an adjacent ratio label set; the adjacent ratio label set includes adjacent importance ratio fields associated with two adjacent indicator fields and generation labels; S500: Based on the adjacent ratio label set and entropy difference set, perform G1 reverse recursive generation of weight vector processing to generate weight vector; S600: Based on the weight vector and quality control index matrix, perform weighted summation to generate a comprehensive risk score sequence; S700: Based on the comprehensive risk score sequence, perform K-means clustering to generate five-level risk labels, and generate five-level risk labels and parameter records; S800, based on five-level risk labels and parameter records, performs GIS-related visualization to generate result data packages, and generates result data packages and risk zoning maps.

2. The method according to claim 1, characterized in that, The process of performing data acquisition and metadata management includes: The evaluation unit boundary data undergoes uniqueness verification and topology consistency checks, including surface feature closure checks, overlap and conflict checks, and hole anomaly checks. Multi-source indicator data undergoes access normalization processing, including data source connection, file parsing, field extraction, and structured encapsulation. The multi-source indicator data includes tabular, raster, and time-series indicator data. Furthermore, spatiotemporal caliber metadata is extracted from the data source header information, product description fields, and caliber dictionary, and the unit caliber identifier undergoes legality verification. A correlation index is established between the spatiotemporal caliber metadata and the multi-source indicator data.

3. The method according to claim 1, characterized in that, The original set of indicators includes: The original indicator set includes evaluation unit identifier field, boundary geometry field, indicator field, value field, time stamp field, spatial positioning field, data source identifier field, version field, and caliber field.

4. The method according to claim 1, characterized in that, The process of performing quality control and caliber alignment includes: Spatiotemporal caliber alignment involves unifying the time stamp field, spatial location field, and caliber field under the constraints of the target time granularity identifier and the target spatial caliber identifier. This spatiotemporal caliber alignment includes time granularity unification, spatial caliber alignment, and unit caliber conversion. Missing data gating determines the missing status and proportion of value fields and triggers hierarchical imputation or deletion actions. Hierarchical imputation generates replacement values ​​for missing items according to three rules: priority of same source and same caliber, temporal proximity, and spatial neighborhood. Anomaly pruning truncates samples in the value fields that exceed the legal range or statistical threshold.

5. The method according to claim 1, characterized in that, The quality control index matrix includes: The quality control index matrix includes evaluation unit identifier, index field and value field, and carries quality markers, which include missing markers, imputation source markers, pruning markers and caliber conversion markers.

6. The method according to claim 1, characterized in that, The process of performing entropy and difference measurement calculations includes: The process includes: standardization of direction, mapping negative indicators to positive indicators, and scaling the value fields to zero to one, with the standardization based on linear scaling of the maximum and minimum values ​​within the same indicator field; entropy calculation, which calculates the entropy value of each indicator field based on the standardized value fields, including ratioization, i.e., proportionalization of the standardized value fields to obtain the ratio field; and difference measurement generation, which generates difference measurement fields based on the entropy values, constructed using a reverse discrete representation of entropy values ​​to minimize difference measurement when the entropy values ​​are close to a uniform distribution.

7. The method according to claim 1, characterized in that, The process of generating the adjacent importance ratio includes: Stable term addition: Preset non-negative stable term values ​​are superimposed onto the difference measurement fields of adjacent preceding and following indicator fields before ratioization; Approach threshold backoff: When the difference between the difference measurement fields of adjacent preceding and following indicator fields is less than a preset approach threshold, the adjacent importance ratio field is assigned a value of one; Upper and lower limit truncation: The adjacent importance ratio field is restricted to between a preset lower limit and a preset upper limit; and monotonic constraint: The sequence of adjacent importance ratio fields is scanned and adjusted to ensure consistency with the sorting relationship of the difference measurement field.

8. The method according to claim 1, characterized in that, The process of generating weight vectors by performing G1 reverse recursion includes: The system employs hierarchical grouping, categorizing indicator fields into three groups: disaster-causing factors, disaster-inducing environment, and disaster-bearing bodies. Within each group, the system sorts the indicator fields in descending order of difference metrics. A G1 reverse recursive approach is used, recursively calculating weight values ​​from the last item forward along a deterministic order, utilizing the adjacent importance ratio field. Non-negative normalization and consistency checks are performed, converting the weight value fields to non-negative values ​​with a sum of one, and executing a monotonic check on the weights within each group. Finally, a weight vector is generated, containing both indicator fields and weight value fields.

9. The method according to claim 1, characterized in that, The process of generating a comprehensive risk score sequence by performing weighted summation includes: The process includes: preparing numerical data with the same caliber; scaling and mapping the value fields in the quality control indicator matrix based on the maximum and minimum value records in the processing log to generate weighted value fields; performing weighted summation on each indicator field within the same evaluation unit identifier and summing them up; and boundary verification to determine whether the comprehensive risk score value field falls within the preset scoring range; and generating a comprehensive risk score sequence, which includes the evaluation unit identifier field and the comprehensive risk score value field.

10. The method according to claim 1, characterized in that, The process of generating five-level risk labels and parameter records by performing K-means clustering includes: K-means clustering is initialized multiple times. Under the premise of a fixed number of five clusters, the clustering iteration is repeatedly performed using different initialization numbers and random number seeds. Silhouette coefficient evaluation is performed by calculating the silhouette coefficient based on the distance between samples within and between clusters, and the clustering result with the largest silhouette coefficient is selected as the optimal result. Five-level risk labels are generated by sorting the cluster centers of the optimal result according to their center values ​​and mapping them to five discrete levels. Five-level risk labels and parameter records are generated. The five-level risk labels include an evaluation unit identifier field and a five-level risk label field. The parameter records include the initialization number, random number seed, cluster center, and silhouette coefficient.