Intelligent government affair information identification processing system and identification method based on big data

By constructing a multi-source integrated smart government information identification system, the problems of inconsistent data standards and broken logical chains in government information identification systems have been solved, realizing the traceability, verifiability, and reusability of government information, and improving the accuracy and compliance of government information identification.

CN121901320AInactive Publication Date: 2026-04-21李宽
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
李宽
Filing Date
2025-12-30
Publication Date
2026-04-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing government identification systems lack the ability to integrate multi-source heterogeneous data, and cannot automatically supplement and verify field gaps, resulting in broken identification chains and discontinuous judgment logic. Furthermore, the lack of a systematic compliance audit and model robustness feedback mechanism can easily lead to distortion of model judgment results or potential biases over time.

Method used

The system constructs a minimum evidence vector generation module, a rule constraint solving module, a state collaboration identification module, a privacy alignment parsing module, a risk routing decision module, and a compliance audit write-back module. By configuring audit indicators such as the maximum mean difference, population stability index, and group identification accuracy difference, it achieves multi-source fusion of government data, logical constraint solving, privacy and security matching, and compliance feedback. It continuously monitors the robustness and fairness of the identification model and automatically triggers model retraining or strategy adjustment.

Benefits of technology

It has enabled the traceability, auditability, and long-term stable operation of government identification results, significantly improving the accuracy, compliance, and intelligence of identification, solving the problems of inconsistent data standards and broken logical chains, and ensuring the traceability and consistency of government information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901320A_ABST
    Figure CN121901320A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent government affair information identification processing system and identification method based on big data, and relates to the field of information identification, and the method comprises the steps: obtaining government affair information and a standardized input set, executing field cutting and unified naming, configuring a minimum evidence vector, generating an evidence cost weight matrix, and recording evidence account book entries; policy caliber field matching is executed, an evidence gap list is formed, and a field-level evidence driving chain is configured. By constructing a minimum evidence vector generation module, a rule constraint solving module, a state collaborative recognition module, a privacy alignment analysis module, a risk routing decision module and a compliance audit write-back module, an intelligent recognition closed-loop system covering the whole process of data acquisition, logic constraint, state tracking, privacy protection, risk decision and compliance audit is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information recognition, and in particular to a smart government information recognition and processing system and recognition method based on big data. Background Technology

[0002] With the deepening of the digital transformation of e-government, government departments at all levels have generated massive amounts of heterogeneous data in the process of administrative approval, public services and policy implementation. This includes structured government affairs data, semi-structured form data and unstructured text information. Traditional government information identification systems rely heavily on rule matching and manual review, lacking the ability to integrate multi-source heterogeneous data. This makes it difficult to achieve intelligent identification, judgment and traceability of government behavior processes. At the same time, due to frequent updates to policy clauses, inconsistent standards, and inconsistent semantics of data fields, existing identification models are prone to problems such as missing fields, biased standards and inconsistent identification conclusions during implementation, affecting the accuracy and transparency of policy implementation.

[0003] However, existing government identification systems have two shortcomings: First, they lack a unified solution mechanism for the logical constraints between input fields, making it impossible to automatically supplement and verify field gaps, resulting in a broken identification chain and discontinuous judgment logic. Second, they have not established a systematic compliance audit and model robustness feedback mechanism, and lack dynamic monitoring of input distribution drift, fairness differences, and population structure shifts in the identification model, which can easily lead to distortion of model judgment results over time or potential biases. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by providing a smart government information identification and processing system and method based on big data. The system aims to achieve multi-source fusion of government data, logical constraint solving, privacy and security matching, and compliance feedback by constructing a minimum evidence vector generation module, a rule constraint solving module, a state collaborative identification module, a privacy alignment parsing module, a risk routing decision module, and a compliance audit write-back module. By configuring audit indicators such as maximum mean difference, population stability index, and group identification accuracy difference, the system can continuously monitor the robustness and fairness of the identification model and automatically trigger model retraining or strategy adjustment. This ensures the traceability, auditability, and long-term stable operation of the identification results, significantly improving the accuracy, compliance, and intelligence level of government information identification.

[0005] Therefore, this application provides a smart government information identification and processing method based on big data, including the following steps:

[0006] Step S100: Obtain government information and standardized input set, perform field pruning and unified naming, configure minimum evidence vector, generate evidence cost weight matrix and record evidence ledger entries, perform policy caliber field matching and form evidence gap list, and configure field-level evidence driving chain.

[0007] Step S200: Configure the policy rule constraint graph, perform constraint verification and three-state result output, trigger micro-evidence supplementation based on the three-state results, generate a supplementation field list, output constraint results, and register the structured rule call path.

[0008] Step S300: Set the administrative event state machine structure, configure the identification task mapping, configure the state synchronization alignment mechanism, form a state graph structure, configure the concurrent state identification mechanism and priority arbitration decision unit, and output the state graph and state snapshot structure.

[0009] It is difficult to find an S400, configure a privacy-preserving entity parsing mechanism, set desensitized field slicing and encrypted signature strategies, perform federated matching and identity field alignment, configure differential privacy budget control, output the alignment result structure, and configure the recognition result aggregation and uniqueness deduplication mechanism.

[0010] Step S500: Configure the counterfactual result inference model and risk cost indicator set, calculate the expected risk cost of the identified processing path, configure the decision function model, configure the high-risk action identification mechanism and route interruption rollback strategy, generate intelligent identification route trajectory, and synchronize the task status diagram and compliance snapshot structure.

[0011] Step S600: Configure the model robustness monitoring and input distribution drift detection mechanism, configure the population structure stability assessment and model fairness difference detection mechanism, configure the model compliance audit data structure and lake warehouse integrated write-back mechanism, and form a four-dimensional conclusion for the identification task.

[0012] In some specific embodiments, step S100 specifically includes:

[0013] Step S100.1: Obtain government information and standardized input set, perform field trimming and unified naming, and configure minimum evidence vector.

[0014] Step S100.2: Generate the evidence cost weight matrix and record evidence ledger entries, perform policy caliber field matching and form an evidence gap list, and configure the field-level evidence-driven chain.

[0015] The minimum evidence vector generation module calculates the evidence cost weight of each field in the minimum evidence vector and generates an evidence cost weight matrix. The evidence cost weight matrix is ​​used to quantify the collection cost and compliance risk of each field. The evidence cost weight modeling unit calculates the field weight score based on the field collection difficulty, field sensitivity level, degree of public data reuse, and historical review value. The weight score adopts a five-level scoring mechanism: the weight of high-sensitivity fields is set to 0.8 or above, the weight of medium-sensitivity fields is set to 0.6 to 0.8, the weight of general fields is set to 0.4 to 0.6, the weight of inferred fields is set to 0.2 to 0.4, and the weight of low-impact fields is set to below 0.2.

[0016] The minimum evidence vector generation module performs policy caliber field matching operations. Through the policy caliber mapping unit, it compares the fields in the minimum evidence vector with the list of caliber fields in the policy rule base, outputs the field consistency judgment result, and forms an evidence gap list. The evidence gap list records the unmet caliber field items and their necessity level. Strongly dependent gap fields are necessary fields for task judgment, and weakly dependent gap fields are auxiliary judgment fields. Each gap field records the field name, caliber field location, and recommended collection path.

[0017] The minimum evidence vector generation module is also used to construct a field-level evidence-driven chain structure. The field-level evidence-driven chain structure consists of field information and its metadata in the minimum evidence vector. The metadata includes the field source path, field collection time, policy version number, identification model fingerprint information, and field integrity verification summary. The identification model fingerprint information includes the identification model name, identification model version number, and identification model key parameter set. The field integrity verification summary is used to verify whether the field content maintains its original consistency. The field-level evidence-driven chain structure is recorded in the evidence ledger.

[0018] In some specific embodiments, step S200 specifically includes:

[0019] Step S200.1: Configure the policy rule constraint graph and perform constraint verification and three-state result output.

[0020] Step S200.2: Trigger micro-evidence supplementation based on the three-state results, generate a list of supplementation fields, output constraint results, and register the structured rule call path.

[0021] In some specific embodiments, step S300 specifically includes:

[0022] Step S300.1: Set the administrative event state machine structure, configure the identification task mapping, configure the state synchronization alignment mechanism, and form a state graph structure.

[0023] Step S300.2: Configure the concurrent state recognition mechanism and priority arbitration decision unit, and output the state graph and state snapshot structure.

[0024] The event state modeling module includes a state arbitration unit, which is used to determine the final dominant state node when there are multiple candidate state nodes. When there is concurrency among the candidate state nodes, the state arbitration unit makes a judgment based on priority rules. The priority rules include: state nodes with earlier event timestamps have higher priority; state nodes with higher field integrity ratios in the state binding evidence set have higher priority; and state nodes with higher binding rule judgment levels have priority.

[0025] After completing the task state graph generation and the determination of the dominant state node, the event state modeling module outputs a state graph structure and a state snapshot structure with state semantics. The state graph structure includes node topology, state transition path, trigger rule number and node binding evidence index. The state snapshot structure records the current state of the task, the last change time, the executed transition history and the set of candidate state paths.

[0026] In some specific embodiments, step S400 specifically includes:

[0027] Step S400.1: Configure the privacy-protected entity parsing mechanism, set the de-identified field slicing and encrypted signature strategy, and perform federated matching and identity field alignment.

[0028] Step S400.2: Configure differential privacy budget control, output the alignment result structure, and configure the recognition result aggregation and uniqueness deduplication mechanism.

[0029] The privacy-preserving entity alignment parsing module includes a privacy budget measurement unit and an alignment ledger registration unit. The privacy budget measurement unit is used to quantify and dynamically limit the differential privacy budget in each federated matching operation. The differential privacy budget parameter is denoted as ε, which characterizes the privacy consumption level of the current identification task. The system sets the privacy budget threshold as: ε max =5.0, when the cumulative privacy budget consumption value ε for the identification task is 5.0. acc Greater than or equal to ε max At that time, the privacy budget metering unit suspends subsequent federal matching operations for the identification task and generates a budget overdraft warning sign.

[0030] The privacy budget metering unit writes the alignment request number, cryptographic signature call record, matching confidence result, alignment assertion type, and ε increment value into the alignment audit ledger. The alignment ledger registration unit then establishes an index binding between this entry and the identification record ledger, forming an auditable alignment tracking record.

[0031] After completing the federated matching process, the privacy-preserving entity alignment parsing module outputs an alignment result structure containing: alignment assertion results, matching source system identifiers, task number mapping relationships, and deduplication identifier sets.

[0032] In some specific embodiments, step S500 specifically includes:

[0033] Step S500.1: Configure the counterfactual outcome inference model and risk cost index set, calculate the expected risk cost of the identified processing path, and configure the decision function model.

[0034] Step S500.2: Configure a high-risk action identification mechanism and a route interruption rollback strategy, generate an intelligent identification route trajectory, and synchronize the task status graph and compliance snapshot structure.

[0035] When the intelligent identification and decision-making routing module detects that the risk cost of the current recommended path exceeds the preset risk tolerance threshold, it triggers a high-risk identification mechanism and executes a route interruption fallback strategy. The risk tolerance threshold is set to: RC max =0.35, when the minimum risk cost value RC i ≥RC max When the intelligent identification and decision-making routing module abandons the automatic processing path and switches to manual review of the processing path, and multiple path risk cost values ​​(RC) exist, the solution is to proceed as follows: i <RC max When the risk cost difference is less than 0.05, the field supplementary evidence processing path is selected, and the field collection instruction is executed by calling the evidence gap list generated by the minimum evidence vector generation module.

[0036] After determining the optimal identification and processing path, the intelligent identification decision-making routing module generates an intelligent identification route trajectory and synchronizes the task status diagram. The intelligent identification decision-making routing module includes a route trajectory generation unit and a compliance archiving unit. The route trajectory generation unit is used to record the identification task number, identification and processing path number, risk cost parameters, recommended path type and final execution result, forming a structured identification route trajectory set.

[0037] In some specific embodiments, step S600 specifically includes:

[0038] Step S600.1: Configure the model robustness monitoring and input distribution drift detection mechanism, and configure the population structure stability assessment and model fairness difference detection mechanism.

[0039] Step S600.2: Configure the model compliance audit data structure and the lake warehouse integrated write-back mechanism, and form a four-dimensional conclusion for the identification task.

[0040] After each identification task is completed, the model robustness and compliance audit feedback module will generate a compliance audit data set by combining the processing results of the identification task, the field-level evidence-driven chain structure, the strategy execution path number, and the model audit indicator results. The compliance audit data set is written into the LakeWarehouse integrated data management platform to form a structured storage form. The fields include the unique identification task number, processing path type, model version number, maximum mean difference value, population stability index value, group identification accuracy difference value, trigger flag, response operation type, and recording time.

[0041] The intelligent government information identification and processing system constructs a four-dimensional closed-loop structure covering the entire identification task process through a model robustness and compliance audit feedback module, specifically including: evidence dimension, strategy dimension, judgment dimension, and response dimension.

[0042] The evidence dimension includes: field-level evidence-driven chain path and evidence integrity summary; the strategy dimension includes: intelligent identification decision routing path and associated strategy number; the judgment dimension includes: identification task judgment result and judgment confidence level; and the response dimension includes: audit indicator trigger record, adjustment behavior and retraining feedback parameters.

[0043] A smart government information identification and processing system based on big data includes the following modules:

[0044] The evidence vector generation module is used to acquire government information and standardized input sets, perform field pruning and unified naming, configure the minimum evidence vector, generate an evidence cost weight matrix and record evidence ledger entries, perform policy caliber field matching and form an evidence gap list, and configure a field-level evidence-driven chain.

[0045] The rule constraint solving module is used to configure the policy rule constraint graph, perform constraint verification and output three-state results, trigger micro-evidence supplementation based on the three-state results, generate a list of supplemented fields, output constraint results, and register the structured rule call path.

[0046] The state collaboration identification module is used to set the state machine structure of administrative events, configure the identification task mapping, configure the state synchronization alignment mechanism, form a state graph structure, configure the concurrent state identification mechanism and priority arbitration decision unit, and output the state graph and state snapshot structure.

[0047] The privacy alignment parsing module is used to configure the privacy-protected entity parsing mechanism, set the de-identified field slicing and encrypted signature strategy, perform federated matching and identity field alignment, configure differential privacy budget control, output the alignment result structure, and configure the recognition result aggregation and uniqueness deduplication mechanism.

[0048] The risk routing decision module is used to configure the counterfactual outcome inference model and risk cost indicator set, calculate the expected risk cost of the identified processing path, configure the decision function model, configure the high-risk action identification mechanism and routing interruption rollback strategy, generate intelligent identification routing trajectory, and synchronize the task status diagram and compliance snapshot structure.

[0049] The compliance audit write-back module is used to configure the model robustness monitoring and input distribution drift detection mechanism, configure the population structure stability assessment and model fairness difference detection mechanism, configure the model compliance audit data structure and lake warehouse integrated write-back mechanism, and generate four-dimensional conclusions for the identification task.

[0050] In summary, this application provides a smart government information identification and processing system and method based on big data. By constructing a minimum evidence vector generation module, a rule constraint solving module, a state collaborative identification module, a privacy alignment parsing module, a risk routing decision module, and a compliance audit write-back module, it forms an intelligent identification closed-loop system covering the entire process of data collection, logical constraints, state tracking, privacy protection, risk decision-making, and compliance audit. Addressing the problems of inconsistent data standards, missing fields, and broken logical chains in existing government identification systems, this invention achieves the trimming, unified mapping, and cost constraint modeling of government information fields through the construction of a minimum evidence vector and an evidence cost weight matrix. It can automatically identify missing fields according to policy rules and trigger micro-evidence supplementation, thereby forming a traceable, verifiable, and reusable field-level evidence-driven chain, effectively improving the quality of data input and the consistency of policy judgments. Attached Figure Description

[0051] Figure 1 This is an overall flowchart of a smart government information identification and processing method based on big data, provided in an embodiment of this application.

[0052] Figure 2 This is an overall framework diagram of a smart government information identification and processing system based on big data, provided in an embodiment of this application. Detailed Implementation

[0053] Please refer to Figure 1 The document illustrates a flowchart of an embodiment of a smart government information identification and processing system and identification method based on big data, in accordance with the present disclosure.

[0054] like Figure 1 As shown, a smart government information identification and processing method based on big data includes the following steps:

[0055] Step S100: Obtain government information and standardized input set, perform field pruning and unified naming, configure minimum evidence vector, generate evidence cost weight matrix and record evidence ledger entries, perform policy caliber field matching and form evidence gap list, and configure field-level evidence driving chain.

[0056] Step S200: Configure the policy rule constraint graph, perform constraint verification and three-state result output, trigger micro-evidence supplementation based on the three-state results, generate a supplementation field list, output constraint results, and register the structured rule call path.

[0057] Step S300: Set the administrative event state machine structure, configure the identification task mapping, configure the state synchronization alignment mechanism, form a state graph structure, configure the concurrent state identification mechanism and priority arbitration decision unit, and output the state graph and state snapshot structure.

[0058] It is difficult to find an S400, configure a privacy-preserving entity parsing mechanism, set desensitized field slicing and encrypted signature strategies, perform federated matching and identity field alignment, configure differential privacy budget control, output the alignment result structure, and configure the recognition result aggregation and uniqueness deduplication mechanism.

[0059] Step S500: Configure the counterfactual result inference model and risk cost indicator set, calculate the expected risk cost of the identified processing path, configure the decision function model, configure the high-risk action identification mechanism and route interruption rollback strategy, generate intelligent identification route trajectory, and synchronize the task status diagram and compliance snapshot structure.

[0060] Step S600: Configure the model robustness monitoring and input distribution drift detection mechanism, configure the population structure stability assessment and model fairness difference detection mechanism, configure the model compliance audit data structure and lake warehouse integrated write-back mechanism, and form a four-dimensional conclusion for the identification task.

[0061] In some specific embodiments, step S100 specifically includes:

[0062] Step S100.1: Obtain government information and standardized input set, perform field trimming and unified naming, and configure minimum evidence vector.

[0063] The big data-driven smart government information identification and processing system includes: a minimum evidence vector generation module, which is used to generate a minimum evidence vector based on government information input data and establish an evidence-driven chain.

[0064] The minimum evidence vector generation module is used to receive structured government data, semi-structured form data and unstructured text data from government service platforms, business processing systems and third-party interaction channels. The data structuring processing unit performs field extraction, format normalization, content decoupling and semantic merging on the structured government data, semi-structured form data and unstructured text data to obtain a standardized input set.

[0065] The minimum evidence vector generation module performs field filtering operations based on the task target variable set. The task target variable set is determined according to the definition of the scope and the acceptance specifications in the policy rule base. The field filtering operation includes calling the field trimming unit, retaining the mandatory fields and reference fields corresponding to the task target variable set from the standardized input set, removing excluded fields, and performing unified mapping on field names with the same semantics but different labels through the unified field mapping table.

[0066] The minimum evidence vector generation module constructs a minimum evidence vector based on the field filtering results. The minimum evidence vector is the minimum set of fields that support the conclusion of government information identification and satisfies the following constraints: field items cannot be redundant, field combinations cannot be reduced, field semantics must be complete and field content can be traced and verified. The minimum evidence vector only contains the necessary fields that can support the judgment conclusion of the current government information identification task.

[0067] Step S100.2: Generate the evidence cost weight matrix and record evidence ledger entries, perform policy caliber field matching and form an evidence gap list, and configure the field-level evidence-driven chain.

[0068] The minimum evidence vector generation module calculates the evidence cost weight of each field in the minimum evidence vector and generates an evidence cost weight matrix. The evidence cost weight matrix is ​​used to quantify the collection cost and compliance risk of each field. The evidence cost weight modeling unit calculates the field weight score based on the field collection difficulty, field sensitivity level, degree of public data reuse, and historical review value. The weight score adopts a five-level scoring mechanism: the weight of high-sensitivity fields is set to 0.8 or above, the weight of medium-sensitivity fields is set to 0.6 to 0.8, the weight of general fields is set to 0.4 to 0.6, the weight of inferred fields is set to 0.2 to 0.4, and the weight of low-impact fields is set to below 0.2.

[0069] The minimum evidence vector generation module performs policy caliber field matching operations. Through the policy caliber mapping unit, it compares the fields in the minimum evidence vector with the list of caliber fields in the policy rule base, outputs the field consistency judgment result, and forms an evidence gap list. The evidence gap list records the unmet caliber field items and their necessity level. Strongly dependent gap fields are necessary fields for task judgment, and weakly dependent gap fields are auxiliary judgment fields. Each gap field records the field name, caliber field location, and recommended collection path.

[0070] The minimum evidence vector generation module is also used to construct a field-level evidence-driven chain structure. The field-level evidence-driven chain structure consists of field information and its metadata in the minimum evidence vector. The metadata includes the field source path, field collection time, policy version number, identification model fingerprint information, and field integrity verification summary. The identification model fingerprint information includes the identification model name, identification model version number, and identification model key parameter set. The field integrity verification summary is used to verify whether the field content maintains its original consistency. The field-level evidence-driven chain structure is recorded in the evidence ledger.

[0071] In some specific embodiments, step S200 specifically includes:

[0072] Step S200.1: Configure the policy rule constraint graph and perform constraint verification and three-state result output.

[0073] The intelligent government information identification and processing system includes a rule constraint solving module. This module is used to perform rule structure modeling and constraint solving operations based on the minimum evidence vector and the policy rule base. The rule constraint solving module includes a rule compilation unit, which is used to parse the clauses, applicable conditions, and field specification requirements in the policy rule base and generate a policy rule constraint graph. The policy rule constraint graph consists of rule nodes and constraint edges. The rule nodes correspond to the task target variables, and the constraint edges describe the inclusion relationship, mutual exclusion relationship, and necessity relationship between fields.

[0074] The rule compilation unit uses a logical constraint-based modeling approach to convert policy rule statements into logical constraint expressions, specifically:

[0075]

[0076] In the formula: R i Let φ represent the i-th policy rule. i (x) represents the set of logical conditions regarding field variable x, y i The rule determination conclusion is represented by all logical constraint expressions connected by Boolean logic to form a unified constraint system. The unified constraint system is mapped to an executable constraint structure for use in the subsequent rule verification process.

[0077] The rule constraint solving module includes a constraint verification unit, which is used to verify the policy rule constraint expression one by one based on the minimum evidence vector. The minimum evidence vector is provided by the minimum evidence vector generation module and includes field content, field source path and policy version information.

[0078] The constraint verification unit compares and verifies the field combination in the minimum evidence vector with the logical expression in the rule constraint diagram, and outputs three state results: satisfied state, violated state, and uncertain state.

[0079] The satisfied state is: all judgment fields in the minimum evidence vector are complete and meet the logical judgment conditions. The violated state is: all judgment fields in the minimum evidence vector exist but do not meet the logical judgment conditions. The uncertain state is: the minimum evidence vector contains fields required for the logical expression that are missing or have empty values.

[0080] The determination status is encoded by a ternary tagging system, and the encoding result is stored in the intermediate cache area for identification and judgment. The intermediate cache area for identification and judgment serves as an intermediate storage unit for storing rule determination status, providing callable results for the risk assessment and identification decision module.

[0081] Step S200.2: Trigger micro-evidence supplementation based on the three-state results, generate a list of supplementation fields, output constraint results, and register the structured rule call path.

[0082] The rule constraint solving module includes a micro-evidence supplementation module, which is used to generate a supplementation field list for rule nodes in an uncertain state during the rule constraint verification process. The supplementation field list records the field name, the policy scope to which the field belongs, the field collection path, the field evidence cost level, and the scope of impact of the supplementation.

[0083] The necessity level threshold for supplementary fields is set to 0.60, and the maximum collection cost threshold is set to 0.75. When the necessity level of a field is ≥0.60 and the collection cost is ≤0.75, the field is included in the supplementary field list. The field collection cost is taken from the weight score of the corresponding field in the evidence cost weight matrix. The evidence cost weight matrix is ​​generated and provided by the minimum evidence vector generation module.

[0084] After completing the constraint solving and three-state determination of all rule nodes, the rule constraint solving module generates a structured rule call path. The structured rule call path records the rule call sequence, rule number, trigger state, field input snapshot, and inference hop count. The rule constraint solving module outputs a set of satisfied rules, a set of violated rules, a set of uncertain rules, and a set of supplementary sampling suggestion fields. All results are bound to a unique identification task number and registered in the identification record ledger. The identification record ledger, as a structured ledger used to register the rule solving process and results, is associated with the evidence ledger to form a three-layer index system of fields, rules, and determinations.

[0085] In some specific embodiments, step S300 specifically includes:

[0086] Step S300.1: Set the administrative event state machine structure, configure the identification task mapping, configure the state synchronization alignment mechanism, and form a state graph structure.

[0087] The intelligent government information identification and processing system includes an event state modeling module. The event state modeling module is used to map the identification results and government event processes into a three-layer administrative event state machine structure with state semantics. The three-layer administrative event state machine structure consists of a set of state nodes, a set of state transition paths, and a set of state binding evidence.

[0088] The event state modeling module calls the rule constraint solution results output by the rule constraint solution module and the minimum evidence vector generated by the minimum evidence vector generation module to determine the initial state node corresponding to each identification task. The state node is used to characterize the stage position of the matter in the government affairs process, the state transition path is used to describe the logical evolution relationship between states, and the state binding evidence set is used to record the field-level evidence-driven chain path that supports the state determination.

[0089] The event state modeling module will uniquely bind the identified task number to the state node and generate a task state graph structure. The task state graph structure is used for subsequent administrative event tracking and dynamic state tracing.

[0090] The event state modeling module includes an event alignment unit, which is used to integrate multi-source time-series information and spatial information from the business processing system, third-party interaction channels and log auditing system to construct an event location vector for state synchronization. The event location vector includes: timestamp parameter, geographic location parameter and main identifier parameter.

[0091] The event alignment unit calls the complex event processing engine to uniformly parse multi-source government event records, extract the timestamp field, spatial location field, and user identity field of each government event, and encapsulate the three fields into an event location vector and bind it to the status node. When there are duplicate task identifiers in multiple systems, the event alignment unit performs the main status node alignment operation according to the event location vector. The main status node alignment priority is executed according to the timestamp priority strategy. The timestamp priority strategy is defined as follows: when the timestamp difference in the event location vector is less than 5 seconds, the spatial location error is less than 30 meters, and the main identifier parameters are consistent, it is determined to be the same main status node.

[0092] The event state modeling module calls the process flow rule table defined in the policy rule library, and marks the current status of the matter, the set of reachable target states and the shortest supplementary evidence path in the task state graph based on the state transition rules. The state graph structure consists of a set of nodes and a set of directed edges, and each directed edge is bound to a state transition rule number and a transition trigger condition.

[0093] Once the transition trigger condition is met, the task state is allowed to transition from the source state node to the target state node.

[0094] When the state transition rule requires supplementary evidence for fields as a prerequisite, the event state modeling module calls the evidence gap list from the minimum evidence vector generation module, filters out the minimum set of fields that support state transition, and forms the shortest supplementary evidence path. The shortest supplementary evidence path is the set of fields that need to be supplemented to transition from the current state to the target state. The module records the field name, field collection path and evidence cost estimate, and outputs a structured supplementary evidence suggestion table.

[0095] Step S300.2: Configure the concurrent state recognition mechanism and priority arbitration decision unit, and output the state graph and state snapshot structure.

[0096] The event state modeling module includes a state arbitration unit, which determines the final dominant state node when multiple candidate state nodes exist. When there is concurrency among candidate state nodes, the state arbitration unit makes a determination based on priority rules. These priority rules include: state nodes with earlier event timestamps have higher priority; state nodes with higher field integrity ratios in their state binding evidence sets have higher priority; and state nodes with higher binding rule judgment levels have priority. The field integrity ratio is calculated as follows:

[0097]

[0098] In the formula: P c N represents the field integrity ratio. a N represents the actual number of fields that exist. t Indicates the total number of fields required for the status.

[0099] When the difference in field integrity ratio is less than 10% and the difference in timestamp is less than 2 seconds, the status node with the smaller strategy number is given priority. The final dominant status node is written into the main node identifier column of the task status diagram for the task flow module and the approval judgment module to call.

[0100] After completing the task state graph generation and the determination of the dominant state node, the event state modeling module outputs a state graph structure and a state snapshot structure with state semantics. The state graph structure includes node topology, state transition path, trigger rule number and node binding evidence index. The state snapshot structure records the current state of the task, the last change time, the executed transition history and the set of candidate state paths.

[0101] In some specific embodiments, step S400 specifically includes:

[0102] Step S400.1: Configure the privacy-protected entity parsing mechanism, set the de-identified field slicing and encrypted signature strategy, and perform federated matching and identity field alignment.

[0103] The intelligent government information identification and processing system includes a privacy-preserving entity alignment and parsing module. This module performs privacy-preserving entity parsing and cross-source deduplication on identity fields involved in the identification task without disclosing personal identity information. The module includes a field slicing unit and an encrypted signature unit. The field slicing unit slices the user identity information field, resident ID number field, and residential address field according to a field rule table. The encrypted signature unit performs homomorphic encrypted hash operations on the field slicing results to generate a structured encrypted signature set.

[0104] The field rule table, provided by the field management module, is used to define the granularity of each type of privacy field and the slice label number. The structured encrypted signature set serves as the input set for the federated matching task.

[0105] The privacy-preserving entity alignment parsing module includes a federated matching engine. This engine performs cross-platform encrypted entity matching and identity field alignment based on a set of structured encrypted signatures, generating entity alignment assertions, alignment confidence assessment values, and field supplementary verification suggestions. The federated matching engine uses a Bloom filter index structure to store encrypted signatures and uses hash cluster distribution and signature co-occurrence frequency as the basis for confidence calculation. The confidence assessment value is defined as follows:

[0106]

[0107] In the formula: C represents the matching confidence score, N match N represents the number of matching hash entries in the cryptographic signature sets at both ends. sig This represents the total number of items in the local encrypted signature set, and ω is the matching weight factor, which is set to 0.95.

[0108] When the matching confidence assessment value is greater than 0.85, the output entity alignment assertion is "alignment confirmed". When the matching confidence assessment value is less than 0.60, the output entity alignment assertion is "mismatch". When the matching confidence assessment value is between 0.60 and 0.85, the output entity alignment assertion is "requires manual review". Based on the evidence gap list, field supplementary evidence suggestions are generated.

[0109] Step S400.2: Configure differential privacy budget control, output the alignment result structure, and configure the recognition result aggregation and uniqueness deduplication mechanism.

[0110] The privacy-preserving entity alignment parsing module includes a privacy budget measurement unit and an alignment ledger registration unit. The privacy budget measurement unit is used to quantify and dynamically limit the differential privacy budget in each federated matching operation. The differential privacy budget parameter is denoted as ε, which characterizes the privacy consumption level of the current identification task. The system sets the privacy budget threshold as: εmax =5.0, when the cumulative privacy budget consumption value ε for the identification task is 5.0. acc Greater than or equal to ε max At that time, the privacy budget metering unit suspends subsequent federal matching operations for the identification task and generates a budget overdraft warning sign.

[0111] The privacy budget metering unit writes the alignment request number, cryptographic signature call record, matching confidence result, alignment assertion type, and ε increment value into the alignment audit ledger. The alignment ledger registration unit then establishes an index binding between this entry and the identification record ledger, forming an auditable alignment tracking record.

[0112] After completing the federated matching process, the privacy-preserving entity alignment parsing module outputs an alignment result structure containing: alignment assertion results, matching source system identifiers, task number mapping relationships, and deduplication identifier sets.

[0113] The deduplication identifier set is used to mark a list of task numbers with unique identity mapping relationships. The recognition result aggregation module calls the deduplication identifier set to perform consistency merging and conflict resolution on multiple recognition task results for the same entity.

[0114] In some specific embodiments, step S500 specifically includes:

[0115] Step S500.1: Configure the counterfactual outcome inference model and risk cost index set, calculate the expected risk cost of the identified processing path, and configure the decision function model.

[0116] The intelligent government information identification and processing system includes an intelligent identification decision-making and routing module. This module is used to establish a counterfactual outcome inference model during the identification task processing and to perform risk assessment and intelligent routing determination of the identification processing path based on a set of risk cost indicators. The intelligent identification decision-making and routing module includes a counterfactual outcome estimation unit and a risk indicator assessment unit. The counterfactual outcome estimation unit is used to construct a causal inference model and a counterfactual consequence model based on historical task samples from the identification record ledger. The identification processing actions include three types: automatic processing path, field supplementary verification processing path, and manual review processing path.

[0117] The counterfactual outcome estimation unit uses a weighted causal forest algorithm to model the consequence variables under different processing paths. The consequence variables include risk consequence indicators such as user complaint rate, administrative reconsideration probability, and government affairs default probability.

[0118] Based on the output of the counterfactual result estimation unit, the risk indicator assessment unit calculates the expected risk cost of the current identification task under different processing paths and constructs a decision function model. The expected risk cost is defined as follows:

[0119]

[0120] Where: RC i R represents the comprehensive risk cost of the i-th identification and processing path. i,j λ represents the predicted consequence value of processing path i under the j-th risk indicator dimension. j This represents the weight coefficient of the j-th risk indicator. The risk indicator dimensions include: the user dissatisfaction complaint rate is weighted at λ1 = 0.35, the government affairs processing misjudgment rate is weighted at λ2 = 0.45, and the administrative reconsideration trigger probability is weighted at λ3 = 0.20.

[0121] The intelligent identification decision routing module selects the identification and processing path with the lowest risk cost as the optimal decision path based on the calculation results, so as to ensure a dynamic balance between risk and efficiency in the identification results.

[0122] Step S500.2: Configure a high-risk action identification mechanism and a route interruption rollback strategy, generate an intelligent identification route trajectory, and synchronize the task status graph and compliance snapshot structure.

[0123] When the intelligent identification and decision-making routing module detects that the risk cost of the current recommended path exceeds the preset risk tolerance threshold, it triggers a high-risk identification mechanism and executes a route interruption fallback strategy. The risk tolerance threshold is set to: RC max =0.35, when the minimum risk cost value RC i ≥RC max When the intelligent identification and decision-making routing module abandons the automatic processing path and switches to manual review of the processing path, and multiple path risk cost values ​​(RC) exist, the solution is to proceed as follows: i <RC max When the risk cost difference is less than 0.05, the field supplementary evidence processing path is selected, and the field collection instruction is executed by calling the evidence gap list generated by the minimum evidence vector generation module.

[0124] The routing interruption and rollback process is uniformly scheduled by the risk policy scheduling unit. All interruption events and parameter records are written into the identification record ledger for subsequent policy optimization and review.

[0125] After determining the optimal identification and processing path, the intelligent identification decision-making routing module generates an intelligent identification route trajectory and synchronizes the task status diagram. The intelligent identification decision-making routing module includes a route trajectory generation unit and a compliance archiving unit. The route trajectory generation unit is used to record the identification task number, identification and processing path number, risk cost parameters, recommended path type and final execution result, forming a structured identification route trajectory set.

[0126] The compliance sealing unit is used to generate a compliance snapshot structure of the task identification process. The compliance snapshot structure is bound to the unique number of the identification task and is linked to the evidence ledger and the identification record ledger to establish a joint index, which supports subsequent compliance audits and risk strategy reviews.

[0127] In some specific embodiments, step S600 specifically includes:

[0128] Step S600.1: Configure the model robustness monitoring and input distribution drift detection mechanism, and configure the population structure stability assessment and model fairness difference detection mechanism.

[0129] The intelligent government information identification and processing system includes a model robustness and compliance audit feedback module. This module monitors the stability of the input distribution and detects model behavior deviations during the execution of identification tasks. The input samples originate from a standardized set of input data. Standardization is achieved through a unified field dictionary and timestamp rules for format alignment. Based on regional and time dimensions, the module compares the current period's input data with historical baseline input data in the feature space, using the maximum mean difference as a measure of input distribution drift. The maximum mean difference is calculated as follows:

[0130] MMD(x,y)=||E[(x)]-E[(y)]|| 2

[0131] In the formula: MMD(x, y) represents the maximum mean difference, used to measure the degree of difference in distribution between the current period's input sample set and the historical baseline input sample set in the feature space; x represents the current period's input sample set; y represents the historical baseline input sample set; (·) represents the kernel function mapping operator, used to map the input sample set to a high-dimensional feature space; (x) represents the feature representation of the current period's input sample set in the kernel function mapping space; (y) represents the feature representation of the historical baseline input sample set in the kernel function mapping space; E[·] represents the expectation operator, used to calculate the mean of the input sample set in the feature space; · represents the Euclidean norm operator, used to calculate the squared distance between two feature mean vectors; δ mmd This represents the maximum mean difference threshold parameter, used to determine whether the input distribution drift is significant. When MMD(x, y) > δ mmd When a significant shift occurs in the input distribution of the identification model, the monitoring period is set to once daily, with a default threshold δ. mmd =0.012.

[0132] The model robustness and compliance audit feedback module includes a population structure stability monitoring unit and a model fairness detection unit. The population structure stability monitoring unit monitors the distribution of demographic characteristics in the input sample to determine whether the sample population structure deviates from the historical baseline distribution. The population structure stability monitoring unit uses the population stability index as the evaluation indicator. The population stability index is calculated as follows:

[0133]

[0134] In the formula: PSI represents the population stability index, which measures the degree of deviation between the current sample population structure and the historical baseline population structure; P i Q represents the proportion of the i-th population characteristic in the current period sample. i This represents the proportion of the i-th demographic characteristic in the historical baseline data, (P) i -Q i This indicates the difference between the current sample percentage and the baseline percentage. The natural logarithm of the ratio between the two is used to measure the direction and intensity of the shift. n represents the total number of population feature groups. When PSI > 0.2, it is determined that the population structure has changed significantly, and the model retraining or structural adjustment mechanism needs to be triggered.

[0135] The model fairness detection unit is used to determine the differences in the model's output performance among different sensitive groups, using the group identification accuracy difference index ΔF. g The quantification is performed, and the calculation is as follows:

[0136] ΔF g =|Acc group1 -Acc group2 |

[0137] In the formula: ΔF g This represents the difference in group recognition accuracy, used to measure the degree of difference in model recognition accuracy between different sensitive groups. |·| represents the absolute value operator, used to calculate the absolute difference in recognition accuracy between two groups. group1 This represents the model's accuracy in identifying the sensitive group 1 sample. Acc group2 This represents the model's accuracy in identifying sensitive group 2 samples. Sensitive groups refer to groups with characteristics subject to policy fairness constraints, such as gender, age group, region, or ethnicity. When ΔF... g When the value is ≥0.15, the model is deemed to have potential unfairness, and policy caliber regression testing and adaptive update of the identification value are required.

[0138] Step S600.2: Configure the model compliance audit data structure and the lake warehouse integrated write-back mechanism, and form a four-dimensional conclusion for the identification task.

[0139] After each identification task is completed, the model robustness and compliance audit feedback module will generate a compliance audit data set by combining the processing results of the identification task, the field-level evidence-driven chain structure, the strategy execution path number, and the model audit indicator results. The compliance audit data set is written into the LakeWarehouse integrated data management platform to form a structured storage form. The fields include the unique identification task number, processing path type, model version number, maximum mean difference value, population stability index value, group identification accuracy difference value, trigger flag, response operation type, and recording time.

[0140] The compliance audit data set, identification record ledger, and evidence ledger are linked together by a unique identification task number to form a joint index for subsequent compliance audits, algorithm risk review, and regulatory tracing.

[0141] The intelligent government information identification and processing system constructs a four-dimensional closed-loop structure covering the entire identification task process through a model robustness and compliance audit feedback module, specifically including: evidence dimension, strategy dimension, judgment dimension, and response dimension.

[0142] The evidence dimension includes: field-level evidence-driven chain path and evidence integrity summary; the strategy dimension includes: intelligent identification decision routing path and associated strategy number; the judgment dimension includes: identification task judgment result and judgment confidence level; and the response dimension includes: audit indicator trigger record, adjustment behavior and retraining feedback parameters.

[0143] The four-dimensional closed-loop structure forms a unique structured record after the recognition task is completed, which is bound to the unique identification task number. This is used to support the traceability, auditability and compliance verification of the entire process from data input to model response, ensuring the stability, transparency and evolution robustness of the system in the long term.

[0144] like Figure 2 As shown, a smart government information identification and processing system based on big data includes the following modules:

[0145] The evidence vector generation module is used to acquire government information and standardized input sets, perform field pruning and unified naming, configure the minimum evidence vector, generate an evidence cost weight matrix and record evidence ledger entries, perform policy caliber field matching and form an evidence gap list, and configure a field-level evidence-driven chain.

[0146] The rule constraint solving module is used to configure the policy rule constraint graph, perform constraint verification and output three-state results, trigger micro-evidence supplementation based on the three-state results, generate a list of supplemented fields, output constraint results, and register the structured rule call path.

[0147] The state collaboration identification module is used to set the state machine structure of administrative events, configure the identification task mapping, configure the state synchronization alignment mechanism, form a state graph structure, configure the concurrent state identification mechanism and priority arbitration decision unit, and output the state graph and state snapshot structure.

[0148] The privacy alignment parsing module is used to configure the privacy-protected entity parsing mechanism, set the de-identified field slicing and encrypted signature strategy, perform federated matching and identity field alignment, configure differential privacy budget control, output the alignment result structure, and configure the recognition result aggregation and uniqueness deduplication mechanism.

[0149] The risk routing decision module is used to configure the counterfactual outcome inference model and risk cost indicator set, calculate the expected risk cost of the identified processing path, configure the decision function model, configure the high-risk action identification mechanism and routing interruption rollback strategy, generate intelligent identification routing trajectory, and synchronize the task status diagram and compliance snapshot structure.

[0150] The compliance audit write-back module is used to configure the model robustness monitoring and input distribution drift detection mechanism, configure the population structure stability assessment and model fairness difference detection mechanism, configure the model compliance audit data structure and lake warehouse integrated write-back mechanism, and generate four-dimensional conclusions for the identification task.

[0151] In practical application, the smart government information identification and processing system firstly uses a minimum evidence vector generation module to perform the access and standardization of government information during the data input stage. This module receives structured government data, semi-structured form data, and unstructured text data from government service platforms, business processing systems, and third-party interactive channels. It then performs field extraction, format normalization, content decoupling, and semantic merging operations to form a standardized input set. The standardized input set is then processed by a field filtering unit based on the task target variable set, which performs field trimming, retaining mandatory and reference fields corresponding to the task target variable set and eliminating excluded fields. Semantic alignment of fields is achieved through a unified field mapping table. The system sets a field filtering threshold of 95%. When the field mapping completion rate of the standardized input set is lower than this threshold, a field re-collection instruction is triggered to ensure the field integrity and traceability of the minimum evidence vector.

[0152] Next, after completing field screening, the minimum evidence vector generation module constructs an evidence cost weight matrix. For each field, a weight score is calculated based on collection difficulty, field sensitivity level, data reuse degree, and historical review value. High-sensitivity fields are weighted at 0.8 to 1.0, medium-sensitivity fields at 0.6 to 0.8, general fields at 0.4 to 0.6, inferred fields at 0.2 to 0.4, and low-impact fields at less than 0.2. The system uses this weight matrix to generate evidence ledger entries and performs policy caliber field matching. The policy caliber mapping unit compares the fields in the minimum evidence vector with the policy rule base's caliber field list, outputting field consistency results. Unmatched fields are recorded in the evidence gap list and marked with their necessity level. When a field's necessity level is not lower than 0.6 and the collection cost is not higher than 0.75, a supplementary collection field list is generated to trigger automatic collection.

[0153] Subsequently, the rule constraint solving module calls the policy rule library during the task execution phase to establish a policy rule constraint graph structure. The rule nodes correspond one-to-one with the task target variables, and the constraint edges are used to describe the inclusion, mutual exclusion, and necessity relationships between fields. The system generates a rule verification structure based on the logical constraint modeling method and performs constraint verification on the minimum evidence vector one by one. The judgment result output is divided into three categories: satisfied state, violated state, and uncertain state. The intermediate buffer of the identification judgment records the state code of each rule to support subsequent risk assessment and identification decision. For rule nodes in an uncertain state, the micro evidence supplementation module automatically generates a list of supplemented fields and updates the identification record ledger to ensure the integrity and continuity of the logical constraint solving process.

[0154] Subsequently, during the execution phase, the event state modeling module maps the identified task results to the administrative event state machine structure. This state machine includes a set of state nodes, a set of state transition paths, and a set of state binding evidence. The event alignment unit integrates multi-source spatiotemporal information from the business processing system, third-party interaction channels, and log auditing system to construct an event location vector and achieve main state node alignment. When the event timestamp difference is less than 5 seconds, the spatial positioning error is less than 30 meters, and the main identifier parameters are consistent, the task state is determined to be the same node. The system updates the task state graph structure according to the process flow rule table in the policy rule base and outputs state snapshot data containing state node topology, state transition paths, and binding evidence indexes for subsequent task flow and approval judgment modules to call.

[0155] Then, the privacy-protected entity alignment and parsing module performs encrypted entity parsing and cross-source identity field alignment during the fusion of multi-source government data. The field slicing unit slices the identity information field, resident certificate field, and residential address field according to the field rule table. The encrypted signature unit performs homomorphic hash operation on the slicing results to generate a set of structured encrypted signatures. The federated matching engine uses a Bloom filter index structure to calculate the matching confidence. When the matching confidence value is higher than 0.85, the alignment is confirmed. When the matching confidence value is lower than 0.60, it is judged as a mismatch. The privacy budget measurement unit suspends subsequent federated matching operations and generates a budget overdraft warning sign to prevent excessive exposure of sensitive information. After the matching is completed, the alignment result structure is generated and registered in the alignment audit ledger to achieve simultaneous protection of data security and identity consistency.

[0156] Finally, the model robustness and compliance audit feedback module performs input distribution drift detection and population structure fairness assessment after the identification task is completed. The system comprehensively monitors based on the maximum mean difference, population stability index, and group identification accuracy difference index. When the MMD value exceeds the threshold of 0.012 or the PSI value is greater than 0.2, the model is automatically retrained. (g) When the value is higher than 0.15, adaptive adjustment of the identification threshold and regression test of policy caliber are performed. All monitoring results, field-level evidence-driven chain structure, identification strategy path and audit indicator values ​​are written into the structured form of the Lake Warehouse integrated platform, forming a four-dimensional closed-loop record of "evidence-strategy-judgment-response" bound to a unique identification task number, realizing full-process traceability, auditability and robustness feedback, and ensuring the reliability and compliance of the system in long-term operation.

Claims

1. A smart government information identification and processing system and identification method based on big data, characterized in that, Includes the following steps: S100: Obtain government information and standardized input sets, perform field trimming and unified naming, configure minimum evidence vector, generate evidence cost weight matrix and record evidence ledger entries, perform policy caliber field matching and form evidence gap list, and configure field-level evidence driving chain. S200: Configure policy rule constraint graph, perform constraint verification and three-state result output, trigger micro-evidence supplementary collection based on three-state results, generate supplementary collection field list, output constraint results, and register structured rule call path; S300: Set the administrative event state machine structure, configure the identification task mapping, configure the state synchronization alignment mechanism, form a state graph structure, configure the concurrent state identification mechanism and priority arbitration decision unit, and output the state graph and state snapshot structure. S400, configure privacy-preserving entity parsing mechanism, set de-identified field slicing and encrypted signature strategy, perform federated matching and identity field alignment, configure differential privacy budget control, output alignment result structure, and configure recognition result aggregation and uniqueness deduplication mechanism; S500, configure the counterfactual outcome inference model and risk cost indicator set, calculate the expected risk cost of the identified processing path, configure the decision function model, configure the high-risk action identification mechanism and route interruption rollback strategy, generate intelligent identification route trajectory, and synchronize the task status diagram and compliance snapshot structure. The S600 configuration includes a robustness monitoring and input distribution drift detection mechanism, a population structure stability assessment and model fairness difference detection mechanism, a model compliance audit data structure and a lake warehouse integrated write-back mechanism, and forms a four-dimensional conclusion for the identification task.

2. The method for identifying and processing smart government information based on big data according to claim 1, characterized in that, S100 specifically includes: S100.1 Obtain government information and standardized input sets, perform field pruning and unified naming, and configure the minimum evidence vector; The big data-driven smart government information identification and processing system includes: a minimum evidence vector generation module, which is used to generate a minimum evidence vector and establish an evidence-driven chain based on government information input data; The minimum evidence vector generation module is used to receive structured government data, semi-structured form data and unstructured text data from government service platforms, business processing systems and third-party interaction channels. The data structuring processing unit performs field extraction, format normalization, content decoupling and semantic merging on the structured government data, semi-structured form data and unstructured text data to obtain a standardized input set. The minimum evidence vector generation module performs a field filtering operation based on the task target variable set. The task target variable set is determined according to the definition of the caliber and the acceptance specifications in the policy rule base. The field filtering operation includes calling the field trimming unit, retaining the mandatory fields and reference fields corresponding to the task target variable set from the standardized input set, removing excluded fields, and performing a unified mapping on the field names with the same semantics but different labels through the field unified mapping table. The minimum evidence vector generation module constructs a minimum evidence vector based on the field filtering results. The minimum evidence vector is the minimum set of fields that supports the conclusion of government information identification and satisfies the following constraints: field items are not redundant, field combinations are not reduced, field semantics must be complete and field content can be traced and verified. The minimum evidence vector only contains the necessary fields that can support the judgment conclusion of the current government information identification task. S100.2, Generate evidence cost weight matrix and record evidence ledger entries, perform policy caliber field matching and form evidence gap list, and configure field-level evidence-driven chain; The minimum evidence vector generation module calculates the evidence cost weight of each field in the minimum evidence vector and generates an evidence cost weight matrix. The evidence cost weight matrix is ​​used to quantify the collection cost and compliance risk of each field. The evidence cost weight modeling unit calculates the field weight score based on the field collection difficulty, field sensitivity level, public data reuse degree and historical review value. The weight score adopts a five-level scoring mechanism: the weight of high-sensitivity fields is set to 0.8 and above, the weight of medium-sensitivity fields is set to 0.6 to 0.8, the weight of general fields is set to 0.4 to 0.6, the weight of inference fields is set to 0.2 to 0.4, and the weight of low-impact fields is set to below 0.

2. The minimum evidence vector generation module performs policy caliber field matching operations. Through the policy caliber mapping unit, it compares the fields in the minimum evidence vector with the list of caliber fields in the policy rule base, outputs the field consistency judgment result, and forms an evidence gap list. The evidence gap list records the unmet caliber field items and their necessity level. Strongly dependent gap fields are necessary fields for task judgment, and weakly dependent gap fields are auxiliary judgment fields. Each gap field records the field name, caliber field location, and recommended collection path. The minimum evidence vector generation module is also used to construct a field-level evidence-driven chain structure. The field-level evidence-driven chain structure consists of field information and its metadata in the minimum evidence vector. The metadata includes the field source path, field collection time, policy version number, identification model fingerprint information, and field integrity verification summary. The identification model fingerprint information includes the identification model name, identification model version number, and identification model key parameter set. The field integrity verification summary is used to verify whether the field content maintains its original consistency. The field-level evidence-driven chain structure is recorded in the evidence ledger.

3. The method for identifying and processing smart government information based on big data according to claim 1, characterized in that, S200 specifically includes: S200.1 Configure the policy rule constraint graph, and perform constraint verification and output the three-state results; The intelligent government information identification and processing system includes a rule constraint solving module. This module is used to perform rule structure modeling and constraint solving operations based on the minimum evidence vector and the policy rule base. The rule constraint solving module includes a rule compilation unit, which is used to parse the clauses, applicable conditions, and field specification requirements in the policy rule base and generate a policy rule constraint graph. The policy rule constraint graph consists of rule nodes and constraint edges. The rule nodes correspond to the task target variables, and the constraint edges describe the inclusion relationship, mutual exclusion relationship, and necessity relationship between fields. The rule compilation unit uses a logical constraint-based modeling approach to convert policy rule statements into logical constraint expressions. The rule constraint solving module includes a constraint verification unit, which is used to verify the policy rule constraint expression one by one based on the minimum evidence vector. The minimum evidence vector is provided by the minimum evidence vector generation module and includes field content, field source path and policy version information. The constraint verification unit compares and verifies the field combination in the minimum evidence vector with the logical expression in the rule constraint diagram, and outputs three state results: satisfied state, violated state, and uncertain state. The satisfied state is: all judgment fields in the minimum evidence vector are complete and meet the logical judgment conditions; the violated state is: all judgment fields in the minimum evidence vector exist but do not meet the logical judgment conditions; the uncertain state is: the minimum evidence vector contains fields required for logical expressions that are missing or have empty values. S200.2 Trigger micro-evidence supplementation based on the three-state results, generate a list of supplementation fields, output constraint results, and register the structured rule call path; The rule constraint solving module includes a micro-evidence supplementation module, which is used to generate a supplementation field list for rule nodes in an uncertain state during the rule constraint verification process. The supplementation field list records the field name, the policy scope to which the field belongs, the field collection path, the field evidence cost level, and the scope of impact of the supplementation. The necessity level threshold for supplementary fields is set to 0.60, and the maximum collection cost threshold is set to 0.

75. When the necessity level of a field is ≥0.60 and the collection cost is ≤0.75, the field is included in the supplementary field list. The field collection cost is taken from the weight score of the corresponding field in the evidence cost weight matrix. The evidence cost weight matrix is ​​generated and provided by the minimum evidence vector generation module. After completing the constraint solving and three-state determination of all rule nodes, the rule constraint solving module generates a structured rule call path. The structured rule call path records the rule call sequence, rule number, trigger state, field input snapshot, and inference hop count. The rule constraint solving module outputs a set of satisfied rules, a set of violated rules, a set of uncertain rules, and a set of supplementary sampling suggestion fields. All results are bound to a unique identification task number and registered in the identification record ledger. The identification record ledger, as a structured ledger used to register the rule solving process and results, is associated with the evidence ledger to form a three-layer index system of fields, rules, and determinations.

4. The method for identifying and processing smart government information based on big data according to claim 1, characterized in that, The S300 specifically includes: S300.

1. Set the administrative event state machine structure, configure the identification task mapping, configure the state synchronization alignment mechanism, and form a state graph structure; The intelligent government information identification and processing system includes an event state modeling module. The event state modeling module is used to map the identification results and government event processes into a three-layer administrative event state machine structure with state semantics. The three-layer administrative event state machine structure consists of a set of state nodes, a set of state transition paths, and a set of state binding evidence. The event state modeling module calls the rule constraint solving result output by the rule constraint solving module and the minimum evidence vector generated by the minimum evidence vector generation module to determine the initial state node corresponding to each identification task. The state node is used to characterize the stage position of the matter in the government affairs process, the state transition path is used to describe the logical evolution relationship between states, and the state binding evidence set is used to record the field-level evidence-driven chain path that supports the state determination. The event state modeling module includes an event alignment unit, which is used to integrate multi-source temporal and spatial information from the business processing system, third-party interaction channels and log auditing system to construct an event location vector for state synchronization. The event location vector includes: timestamp parameter, geographic location parameter and main identifier parameter. The event alignment unit calls the complex event processing engine to uniformly parse multi-source government event records, extracts the timestamp field, spatial location field, and user identity field of each government event, combines and encapsulates the three fields into an event location vector and binds it to the status node. When there are duplicate task identifiers in multiple systems, the event alignment unit performs the main status node alignment operation according to the event location vector. The main status node alignment priority is executed according to the timestamp priority strategy. The timestamp priority strategy is defined as follows: when the timestamp difference in the event location vector is less than 5 seconds, the spatial location error is less than 30 meters, and the main identifier parameters are consistent, it is determined to be the same main status node. The event state modeling module calls the process flow rule table defined in the policy rule library, and marks the current status of the matter, the set of reachable target states and the shortest supplementary evidence path in the task state graph based on the state transition rules. The state graph structure consists of a set of nodes and a set of directed edges. Each directed edge is bound to the state transition rule number and the transition trigger condition. Once the transition triggering condition is met, the task state is allowed to transition from the source state node to the target state node; When the state transition rule requires supplementary evidence for fields as a prerequisite, the event state modeling module calls the evidence gap list from the minimum evidence vector generation module, filters out the minimum set of fields that support state transition, forms the shortest supplementary evidence path, and the shortest supplementary evidence path is the set of fields that need to be supplemented to transition from the current state to the target state. It records the field name, field collection path and evidence cost estimate, and outputs a structured supplementary evidence suggestion table. S300.2 Configure concurrent state recognition mechanism and priority arbitration decision unit, and output state graph and state snapshot structure; The event state modeling module includes a state arbitration unit, which is used to determine the final dominant state node when there are multiple candidate state nodes. When there is concurrency among the candidate state nodes, the state arbitration unit makes a judgment based on priority rules. The priority rules include: state nodes with earlier event timestamps have higher priority; state nodes with higher field integrity ratios in the state binding evidence set have higher priority; and state nodes with higher binding rule judgment levels have priority. After completing the task state graph generation and the determination of the dominant state node, the event state modeling module outputs a state graph structure and a state snapshot structure with state semantics. The state graph structure includes node topology, state transition path, trigger rule number and node binding evidence index. The state snapshot structure records the current state of the task, the last change time, the executed transition history and the set of candidate state paths.

5. The method for identifying and processing smart government information based on big data according to claim 1, characterized in that, The S400 specifically includes: S400.1 Configure a privacy-preserving entity parsing mechanism, set de-identified field slicing and encrypted signature strategies, and perform federated matching and identity field alignment; The intelligent government information identification and processing system includes a privacy-preserving entity alignment and parsing module. This module performs privacy-preserving entity parsing and cross-source deduplication on identity fields involved in the identification task without disclosing personal identity information. The privacy-preserving entity alignment and parsing module includes a field slicing unit and an encrypted signature unit. The field slicing unit performs field-level slicing processing on the user identity information field, resident ID number field, and residential address field according to a field rule table. The encrypted signature unit performs homomorphic encrypted hash operations on the field slicing results to generate a set of structured encrypted signatures. The field rule table is provided by the field management module and is used to define the granularity of each type of privacy field and the slice label number. The structured encrypted signature set serves as the input set for the federated matching task. The privacy-preserving entity alignment parsing module includes a federated matching engine, which is used to perform cross-platform encrypted entity matching and identity field alignment based on a set of structured encrypted signatures, and generate entity alignment assertions, alignment confidence assessment values ​​and field supplementation suggestions. The federated matching engine uses a Bloom filter index structure to store encrypted signatures and uses hash cluster distribution and signature co-occurrence frequency as the basis for confidence calculation. When the matching confidence assessment value is greater than 0.85, the output entity alignment assertion is "alignment confirmed"; when the matching confidence assessment value is less than 0.60, the output entity alignment assertion is "mismatch"; when the matching confidence assessment value is between 0.60 and 0.85, the output entity alignment assertion is "requires manual review"; and field supplementary evidence suggestions are generated in conjunction with the evidence gap list. S400.2 Configure differential privacy budget control, output aligned result structure, and configure recognition result aggregation and uniqueness deduplication mechanism; The privacy-preserving entity alignment parsing module includes a privacy budget measurement unit and an alignment ledger registration unit. The privacy budget measurement unit is used to quantify and dynamically limit the differential privacy budget in each federated matching operation. The differential privacy budget parameter is denoted as ε, which characterizes the privacy consumption level of the current identification task. The system sets the privacy budget threshold as: ε max =5.0, when the cumulative privacy budget consumption value ε for the identification task is 5.

0. acc Greater than or equal to ε max At this time, the privacy budget metering unit suspends subsequent federal matching operations for the identification task and generates a budget overdraft warning sign; The privacy budget measurement unit writes the alignment request number, cryptographic signature call record, matching confidence result, alignment assertion type and ε increment value into the alignment audit ledger. The alignment ledger registration unit establishes an index binding between this entry and the identification record ledger to form an auditable alignment tracking record. After completing the federated matching process, the privacy-preserving entity alignment parsing module outputs an alignment result structure containing: alignment assertion results, matching source system identifiers, task number mapping relationships, and deduplication identifier sets.

6. The method for identifying and processing smart government information based on big data according to claim 1, characterized in that, The S500 specifically includes: S500.1 Configure the counterfactual outcome inference model and risk cost index set, calculate the expected risk cost of the identified processing path, and configure the decision function model; The intelligent government information identification and processing system includes an intelligent identification decision-making and routing module. This module is used to establish a counterfactual result inference model during the identification task processing and to perform risk assessment and intelligent routing determination of the identification processing path based on a set of risk cost indicators. The intelligent identification decision-making and routing module includes a counterfactual result estimation unit and a risk indicator assessment unit. The counterfactual result estimation unit is used to construct a causal inference model and a counterfactual consequence model based on historical task samples from the identification record ledger. The identification processing actions include three types: automatic processing path, field supplementary evidence processing path, and manual review processing path. The counterfactual outcome estimation unit uses a weighted causal forest algorithm to model the consequence variables under different processing paths. The consequence variables include risk consequence indicators such as user complaint rate, administrative reconsideration probability, and government affairs breach probability. Based on the output of the counterfactual result estimation unit, the risk indicator assessment unit calculates the expected risk cost value of the current identification task under different processing paths and constructs a decision function model. S500.2 Configure a high-risk action identification mechanism and a route interruption rollback strategy to generate intelligent identification route trajectories and synchronize task status graphs and compliance snapshot structures; When the intelligent identification and decision-making routing module detects that the risk cost of the current recommended path exceeds the preset risk tolerance threshold, it triggers a high-risk identification mechanism and executes a route interruption fallback strategy. The risk tolerance threshold is set to: RC max =0.35, when the minimum risk cost value RC i ≥RC max When the intelligent identification and decision-making routing module abandons the automatic processing path and switches to manual review of the processing path, and multiple path risk cost values ​​(RC) exist, the solution is to proceed as follows: i <RC max When the risk cost difference is less than 0.05, the field supplementary evidence processing path is selected, and the field collection instruction is executed by calling the evidence gap list generated by the minimum evidence vector generation module. After determining the optimal identification and processing path, the intelligent identification decision-making routing module generates an intelligent identification route trajectory and synchronizes the task status diagram. The intelligent identification decision-making routing module includes a route trajectory generation unit and a compliance archiving unit. The route trajectory generation unit is used to record the identification task number, identification and processing path number, risk cost parameters, recommended path type and final execution result, forming a structured identification route trajectory set.

7. The method for identifying and processing smart government information based on big data according to claim 1, characterized in that, The S600 specifically includes: S600.1 Configure a model robustness monitoring and input distribution drift detection mechanism, and configure a population structure stability assessment and model fairness difference detection mechanism; The intelligent government information identification and processing system includes a model robustness and compliance audit feedback module. This module is used to monitor the stability of the input distribution and detect model behavior deviations during the execution of the identification task. The input samples are derived from a standardized set of input data. The standardization process achieves format alignment through a unified field dictionary and timestamp rules. Based on regional and time dimensions, the model robustness and compliance audit feedback module compares the current period's input data with historical baseline input data in the feature space and uses the maximum mean difference index as a measure of input distribution drift. The model robustness and compliance audit feedback module includes a population structure stability monitoring unit and a model fairness detection unit. The population structure stability monitoring unit is used to monitor the distribution of demographic characteristics in the input sample and determine whether the sample population structure deviates from the historical baseline distribution. The population structure stability monitoring unit uses the population stability index as the evaluation indicator. The model fairness detection unit is used to determine the differences in the model's output performance among different sensitive groups, using the group identification accuracy difference index ΔF. g Quantify it. S600.2 Configure the model compliance audit data structure and lake warehouse integrated write-back mechanism, and form a four-dimensional conclusion for the identification task; After each identification task is completed, the model robustness and compliance audit feedback module will generate a compliance audit data set by combining the processing results of the identification task, the field-level evidence-driven chain structure, the strategy execution path number, and the model audit indicator results. The compliance audit data set will be written into the Lake Warehouse integrated data management platform to form a structured storage form. The fields include the unique identification task number, processing path type, model version number, maximum mean difference value, population stability index value, group identification accuracy difference value, trigger flag, response operation type, and recording time. The intelligent government information identification and processing system constructs a four-dimensional closed-loop structure covering the entire identification task process through a model robustness and compliance audit feedback module, specifically including: evidence dimension, strategy dimension, judgment dimension and response dimension; The evidence dimension includes: field-level evidence-driven chain path and evidence integrity summary; the strategy dimension includes: intelligent identification decision routing path and associated strategy number; the judgment dimension includes: identification task judgment result and judgment confidence level; and the response dimension includes: audit indicator trigger record, adjustment behavior and retraining feedback parameters.

8. A smart government information identification and processing system based on big data, characterized in that, Includes the following modules: The evidence vector generation module is used to acquire government information and standardized input sets, perform field pruning and unified naming, configure the minimum evidence vector, generate an evidence cost weight matrix and record evidence ledger entries, perform policy caliber field matching and form an evidence gap list, and configure a field-level evidence driving chain. The rule constraint solving module is used to configure the policy rule constraint graph, perform constraint verification and output three-state results, trigger micro-evidence supplementation based on the three-state results, generate a list of supplemented fields, output constraint results, and register the structured rule call path; The state collaboration identification module is used to set the state machine structure of administrative events, configure the identification task mapping, configure the state synchronization alignment mechanism, form a state graph structure, configure the concurrent state identification mechanism and priority arbitration decision unit, and output the state graph and state snapshot structure. The privacy alignment parsing module is used to configure the privacy-protected entity parsing mechanism, set the de-identified field slicing and encrypted signature strategy, perform federated matching and identity field alignment, configure differential privacy budget control, output the alignment result structure, and configure the recognition result aggregation and uniqueness deduplication mechanism. The risk routing decision module is used to configure the counterfactual outcome inference model and risk cost indicator set, calculate the expected risk cost of the identified processing path, configure the decision function model, configure the high-risk action identification mechanism and routing interruption rollback strategy, generate intelligent identification routing trajectory, and synchronize the task status diagram and compliance snapshot structure. The compliance audit write-back module is used to configure the model robustness monitoring and input distribution drift detection mechanism, configure the population structure stability assessment and model fairness difference detection mechanism, configure the model compliance audit data structure and lake warehouse integrated write-back mechanism, and generate four-dimensional conclusions for the identification task.