Intelligent auxiliary decision-making method and system for government examination and approval based on multi-source data fusion
By performing structured parsing and semantic mapping on multi-source data from the government approval system, establishing a traceability chain and resolving conflicts, and constructing a knowledge graph for path reasoning, the problem of data inconsistency in government approval was solved, and the credible measurement of data and the reliability of decision-making were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING YIZHUANG INTELLIGENT CITY RES INST GRP CO LTD
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-22
Smart Images

Figure CN121707516B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of government data processing technology, and in particular to an intelligent auxiliary decision-making method and system for government approval based on multi-source data fusion. Background Technology
[0002] Intelligent government approval is a key direction for improving the efficiency of government services. Its core lies in how to effectively integrate multi-source data to achieve intelligent auxiliary decision-making in the approval process.
[0003] Existing government approval systems lack sufficient semantic understanding and alignment capabilities for multi-source heterogeneous data. Differences in data formats, encoding standards, and semantic expressions exist between different business systems and data sources, leading to inconsistent representations of the same entity across different systems and hindering effective information integration and comparison. Current technologies lack quantitative assessment mechanisms for data credibility, making it impossible to distinguish the reliability of data from different sources. When faced with data conflicts, reasonable judgments are difficult to make, potentially leading to approval decisions based on unreliable information and increasing approval risks. Existing government approval systems largely rely on form-based processing, lacking the ability to model and reason about complex relationships between business entities. This makes it difficult to discover hidden risk correlations and logical contradictions, and prevents the provision of intelligent approval suggestions based on knowledge reasoning. The approval process still heavily depends on human experience and judgment. Summary of the Invention
[0004] This invention provides an intelligent auxiliary decision-making method and system for government approval based on multi-source data fusion, which can solve the problems in the prior art.
[0005] A first aspect of this invention provides an intelligent auxiliary decision-making method for government approval based on multi-source data fusion, comprising:
[0006] The data from various sources in the multi-source heterogeneous data set of government approval business are structured and parsed, and similar entity objects in different source data are mapped to a unified semantic space to obtain standardized data after semantic alignment.
[0007] A traceability chain is established. By calculating the consistency frequency and node verification pass rate of each attribute value in the standardized data, dynamic trustworthy quantification values are assigned to each attribute value, resulting in an enhanced data representation with traceability tags and trustworthy quantification values.
[0008] The attribute values of the same approval item are compared in the enhanced data representation to identify attribute value conflicts between different source data. Based on the credible quantification value of each conflicting attribute value, the integrity of the traceability chain, and the semantic dependency between attributes, multiple attribute values are fused to obtain a fused data entity with conflict resolution.
[0009] Using the fused data entities as nodes and the business relationships in government approval items as edges, a government approval knowledge graph is constructed. Based on graph reasoning rules, path reasoning is performed on the government approval knowledge graph to identify the incompleteness of application materials and the risk association of related entities, and a set of reasoning results is obtained.
[0010] Based on the set of reasoning results, combined with approval business rules and quantitative evaluation criteria, the risk level of the application is quantified and a decision path is recommended to generate auxiliary decision-making results.
[0011] The data from various sources in the multi-source heterogeneous dataset of government approval processes are structured and parsed. Similar entity objects from different sources are mapped to a unified semantic space, resulting in standardized data with semantic alignment, including:
[0012] Semantic features are extracted from each source data in the multi-source heterogeneous data set. The statistical distribution characteristics of each field in each source data are analyzed. Each field is represented as a composite feature vector according to the statistical distribution characteristics. The entity objects in each source data are represented as entity feature matrices according to the corresponding composite feature vectors.
[0013] In the standard semantic space, standard composite feature vectors are generated according to the attribute definition specifications of each standard entity type, and all standard composite feature vectors of the standard entity type are combined into a standard entity type prototype matrix.
[0014] Calculate the matrix field similarity between the entity feature matrix of each source data and the standard entity type prototype matrix, and aggregate the matrix field similarity in the entity feature matrix to obtain the entity-level semantic mapping score;
[0015] The mapping relationship between each source data entity object and the standard entity type is determined based on the entity-level semantic mapping score. The renaming relationship between each source data field and the standard field name is determined based on the matrix field similarity. Based on the mapping relationship and the renaming relationship, the entity type of each source data is converted and the field name is standardized to obtain semantically aligned standardized data.
[0016] A traceability chain is established. By calculating the consistency frequency and node verification pass rate of each attribute value in the standardized data, dynamic trust metrics are assigned to each attribute value, resulting in an enhanced data representation with traceability tags and trust metrics, including:
[0017] Extract each attribute value and historical correction record from the standardized data and organize them into the traceability chain in chronological order;
[0018] For each attribute value in the standardized data, retrieve historical attribute values that are the same as the current attribute value from the set of attribute values of completed approval items, and calculate the consistency frequency between the approval results corresponding to the historical attribute values and the expected approval results; for each transmission path node in the traceability chain, extract the verification records of the current transmission path node performing verification operations on the data, and calculate the node verification pass rate by the ratio between the number of successful verifications in the verification records and the total number of verifications.
[0019] The consistency frequency and the path verification pass rate are weighted and summed, and the weighted summation result is attenuated according to the number of corrections in the historical correction records to obtain a dynamic trust metric value. The dynamic trust metric value is associated with the corresponding traceability chain to the corresponding attribute value in the standardized data to obtain an enhanced data representation with traceability tags and trust metric values.
[0020] The enhanced data representation is used to compare attribute values for the same approval item, and the identification of attribute value conflicts between different source data includes:
[0021] For the same approval item, an attribute value set of the same attribute from different source data is extracted from the enhanced data representation. Each attribute value in the attribute value set is compared pairwise. When there is a numerical difference or semantic difference, it is marked as a candidate conflicting attribute value pair.
[0022] Extract inter-attribute dependency rules related to the candidate conflicting attribute value pairs from a predefined attribute semantic dependency rule base. Extract the attribute values of the associated attributes from the enhanced data representation according to the inter-attribute dependency rules. Determine whether the attribute values in the candidate conflicting attribute value pairs and the attribute values of the associated attributes satisfy the value constraint relationship. If the value constraint relationship is not satisfied, confirm the candidate conflicting attribute value pairs as real conflicting attribute value pairs.
[0023] For each attribute value in the real conflict attribute value pair, the corresponding traceability chain is extracted, and the traceability integrity of the data source identifier and the continuity integrity of the transmission path node sequence in the traceability chain are calculated. The traceability integrity and continuity integrity are weighted and summed to obtain the traceability chain integrity score. The traceability chain integrity score and the credible quantification value of the corresponding attribute value are associated with the real conflict attribute value pair to obtain the attribute value conflict.
[0024] Based on the credible quantification values of each conflicting attribute value, the completeness of the tracing chain, and the semantic dependencies between attributes, multiple attribute values are fused to obtain a conflict-resolved fused data entity, including:
[0025] The credible quantification value associated with each attribute value in the real conflict attribute value pair and the traceability chain integrity score are extracted and weighted to obtain a comprehensive credibility score. The attribute values in the real conflict attribute value pair are sorted according to the comprehensive credibility score, and the attribute value with the highest comprehensive credibility score is selected as the initial fusion attribute value.
[0026] According to the attribute dependency rules, the attribute values of the corresponding associated attributes are extracted from the enhanced data representation. The degree of conformity between the initial fusion attribute value and the attribute values of each associated attribute is calculated to satisfy the attribute dependency rules. The degree of conformity corresponding to all attribute dependency rules is weighted and aggregated to obtain the semantic consistency score.
[0027] The overall credibility score of the initial fusion attribute value is adjusted based on the semantic consistency score and the semantic dependency strength defined by the attribute dependency rules to obtain the final fusion attribute value;
[0028] The final fused attribute value replaces the corresponding real conflict attribute value pair in the enhanced data representation, and the traceability chain and trusted quantification associated with the final fused attribute value are retained in the replaced enhanced data representation to obtain the fused data entity after conflict resolution.
[0029] Based on graph reasoning rules, path reasoning is performed on the government approval knowledge graph to identify incomplete or missing application materials and risk associations with related entities, resulting in a set of reasoning results including:
[0030] Extract multi-level material dependency reasoning rules and subject association propagation reasoning rules from a predefined graph reasoning rule base, and perform multi-path parallel traversal in the government approval knowledge graph starting from the approval item identification node;
[0031] Calculate the structural similarity between each traversal path and each path pattern in the multi-level material dependency reasoning rules, select the material necessity weight corresponding to the path pattern with the highest structural similarity as the weight value of the traversal path, aggregate the weight values of multiple traversal paths that reach the same material identifier node to obtain the aggregate necessity score of the material identifier node, and mark the material identifier node as a material missing item when the aggregate necessity score is lower than the preset material missing judgment threshold;
[0032] Extract approval entity identifier nodes with risk identification attributes from the government approval knowledge graph as risk source nodes. Perform breadth-first traversal along the business relationship edges starting from the risk source nodes. During the traversal, calculate the risk propagation score of the current node according to the entity association propagation reasoning rules. Extract approval entity identifier nodes with risk propagation scores higher than the preset risk association judgment threshold as risk-related entities and record the risk propagation path.
[0033] The missing material items and their corresponding aggregation necessity scores, the risk-related entities, and the risk propagation paths are encapsulated into a set of inference results.
[0034] Based on the aforementioned set of reasoning results, and in conjunction with approval business rules and quantitative assessment criteria, the risk level of the application is quantified and a decision-making path is recommended, generating auxiliary decision-making results including:
[0035] Extract the aggregation necessity score of missing materials and the risk propagation score of the risk-associated subject from the set of inference results; obtain the material type importance coefficient and calculate the material integrity risk score by weighting it with the aggregation necessity score; obtain the subject risk type weight coefficient and calculate the subject association risk score by weighting it with the risk propagation score.
[0036] The material integrity risk score and the subject association risk score are fused from multiple dimensions to obtain a comprehensive risk assessment score. The comprehensive risk assessment score is then mapped to the corresponding risk level identifier according to a preset risk level classification rule.
[0037] Based on the risk level identifier, the missing materials are prioritized according to the aggregated necessity score and a supplementary material list is generated. The risk propagation path in the inference result set is decomposed, and the business relationship edge type sequence is extracted. Based on the business relationship edge type sequence, a targeted risk verification process suggestion is generated. Based on the risk level identifier, a predefined approval process processing strategy is matched to generate an approval process optimization scheme.
[0038] The comprehensive risk assessment score, the risk level identifier, the list of supplementary materials, the risk verification process suggestions, and the approval process optimization plan are packaged into the auxiliary decision-making results.
[0039] A second aspect of this invention provides an intelligent auxiliary decision-making system for government approval based on multi-source data fusion, comprising:
[0040] The first unit is used to perform structured parsing of the source data in the multi-source heterogeneous data set of government approval business, and to map the same type of entity objects in different source data to a unified semantic space to obtain standardized data after semantic alignment.
[0041] The second unit is used to establish a traceability chain. By calculating the consistency frequency and node verification pass rate of each attribute value in the standardized data, dynamic trust quantification values are assigned to each attribute value, resulting in an enhanced data representation with traceability tags and trust quantification values.
[0042] The third unit is used to compare the attribute values of the same approval item in the enhanced data representation, identify attribute value conflicts between different source data, and fuse multiple attribute values based on the credible quantification value of each conflicting attribute value, the integrity of the traceability chain, and the semantic dependency relationship between attributes to obtain a fused data entity after conflict resolution.
[0043] The fourth unit is used to construct a government approval knowledge graph by taking the fused data entities as nodes and the business relationships in government approval items as edges, and to perform path reasoning on the government approval knowledge graph based on graph reasoning rules to identify the incompleteness of application materials and the risk association of related entities, and to obtain a set of reasoning results.
[0044] The fifth unit is used to quantify the risk level and recommend decision-making paths for the application items based on the set of reasoning results, combined with the approval business rules and quantitative evaluation criteria, and to generate auxiliary decision-making results.
[0045] A third aspect of the present invention,
[0046] An electronic device is provided, comprising:
[0047] processor;
[0048] Memory used to store processor-executable instructions;
[0049] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0050] Fourth aspect of the embodiments of the present invention,
[0051] A computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0052] The beneficial effects of this application are as follows:
[0053] By employing structured parsing and unified semantic space mapping, the heterogeneity of government data from different sources in terms of format, structure, and semantic expression is resolved, enabling data to be processed and analyzed under unified standards. A traceability chain is established, and dynamic credibility is calculated using consistency frequency and node verification pass rate, making data credibility quantifiable and traceable, thus improving the transparency of data processing and the reliability of approval decisions. Based on quantifiable credibility values, traceability chain integrity, and semantic dependencies, attribute value conflicts are resolved, achieving intelligent fusion of multi-source data and solving the decision-making difficulties caused by data inconsistency in traditional government approval processes. Attached Figure Description
[0054] Figure 1This is a flowchart illustrating the intelligent auxiliary decision-making method for government approval based on multi-source data fusion, as described in an embodiment of the present invention.
[0055] Figure 2 This is a schematic diagram of the path reasoning process for the knowledge graph of government approval. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0057] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0058] Figure 1 This is a flowchart illustrating an intelligent auxiliary decision-making method for government approval based on multi-source data fusion, as described in an embodiment of the present invention. Figure 1 As shown, the method includes:
[0059] The data from various sources in the multi-source heterogeneous data set of government approval business are structured and parsed, and similar entity objects in different source data are mapped to a unified semantic space to obtain standardized data after semantic alignment.
[0060] A traceability chain is established. By calculating the consistency frequency and node verification pass rate of each attribute value in the standardized data, dynamic trustworthy quantification values are assigned to each attribute value, resulting in an enhanced data representation with traceability tags and trustworthy quantification values.
[0061] The attribute values of the same approval item are compared in the enhanced data representation to identify attribute value conflicts between different source data. Based on the credible quantification value of each conflicting attribute value, the integrity of the traceability chain, and the semantic dependency between attributes, multiple attribute values are fused to obtain a fused data entity with conflict resolution.
[0062] Using the fused data entities as nodes and the business relationships in government approval items as edges, a government approval knowledge graph is constructed. Based on graph reasoning rules, path reasoning is performed on the government approval knowledge graph to identify the incompleteness of application materials and the risk association of related entities, and a set of reasoning results is obtained.
[0063] Based on the set of reasoning results, combined with approval business rules and quantitative evaluation criteria, the risk level of the application is quantified and a decision path is recommended to generate auxiliary decision-making results.
[0064] In one optional implementation, the data from each source in the multi-source heterogeneous data set of government approval business is structured and parsed to map similar entity objects from different source data to a unified semantic space, resulting in semantically aligned standardized data, including:
[0065] Semantic features are extracted from each source data in the multi-source heterogeneous data set. The statistical distribution characteristics of each field in each source data are analyzed. Each field is represented as a composite feature vector according to the statistical distribution characteristics. The entity objects in each source data are represented as entity feature matrices according to the corresponding composite feature vectors.
[0066] In the standard semantic space, standard composite feature vectors are generated according to the attribute definition specifications of each standard entity type, and all standard composite feature vectors of the standard entity type are combined into a standard entity type prototype matrix.
[0067] Calculate the matrix field similarity between the entity feature matrix of each source data and the standard entity type prototype matrix, and aggregate the matrix field similarity in the entity feature matrix to obtain the entity-level semantic mapping score;
[0068] The mapping relationship between each source data entity object and the standard entity type is determined based on the entity-level semantic mapping score. The renaming relationship between each source data field and the standard field name is determined based on the matrix field similarity. Based on the mapping relationship and the renaming relationship, the entity type of each source data is converted and the field name is standardized to obtain semantically aligned standardized data.
[0069] Government data typically originates from different systems across various departments and comes in diverse formats, including structured database tables, semi-structured XML documents, and unstructured PDF files. For this heterogeneous data, basic structuring transformation is first achieved through data preprocessing. For database tables, the table structure and data content are directly extracted. For semi-structured documents such as XML, node information is extracted using a parser and converted into a table format. For unstructured documents such as PDFs, text analysis techniques are used to extract key information and store it in a structured format.
[0070] After initial structuring, semantic features are extracted from each source data, and the statistical distribution characteristics of each field in each data source are analyzed, including features such as data type, value range, numerical distribution, text length, and formatting patterns. For example, the "Applicant Name" field is represented as a fixed-length Chinese string without numbers or special symbols; while the "Contact Number" field is a numeric string conforming to a specific format.
[0071] Based on these statistical distribution characteristics, each field is represented as a composite feature vector. This composite feature vector contains multiple dimensions: data type characteristics (e.g., character, integer, date), format characteristics (e.g., length distribution, character type), value characteristics (e.g., value range, enumerated value set), and business characteristics (e.g., frequency of occurrence, related fields). For example, a phone number field can be represented as: [character type, 11-13 digits long, all numbers, distribution conforms to mobile phone number rules, high correlation with contact person]. The composite feature vectors corresponding to all fields in each entity object are combined to form an entity feature matrix. For example, the feature matrix of an "applicant" entity may contain feature vectors for multiple fields such as name, ID number, and contact information.
[0072] In the standard semantic space, standard composite feature vectors are generated based on the standard entity types and their attribute definition specifications predefined in the field of government approval. For example, the standard "natural person applicant" entity type includes standard attributes such as "name", "ID number" and "contact number", each of which has its own standard definition and constraints.
[0073] Based on these standard definitions and constraints, a standard composite feature vector is generated for each standard attribute, and all standard composite feature vectors of the same standard entity type are combined into a standard entity type prototype matrix. For example, the standard prototype matrix for "natural person applicant" contains the feature vector set of each standard attribute.
[0074] A multi-dimensional matching algorithm is used to calculate the matrix field similarity between the entity feature matrix of each source data and the standard entity type prototype matrix. Different similarity measurement methods are used for different feature dimensions, such as exact matching for data types, pattern similarity for text formats, and distribution similarity for value distributions.
[0075] After the field-level similarity calculation is completed, the field similarity is aggregated into an entity-level semantic mapping score through a weighted aggregation method. The weights can be dynamically adjusted according to the importance of the fields. For example, in the "natural person applicant" entity, the weight of "ID number" is higher than that of "contact address".
[0076] Based on entity-level semantic mapping scores, the mapping relationship between each source data entity object and the standard entity type is determined. The mapping rules can set thresholds. When the mapping score exceeds a specific threshold (such as 0.85), the source data entity is mapped to the corresponding standard entity type.
[0077] At the same time, the renaming relationship between each source data field and the standard field name is determined based on the matrix field similarity. For example, the "Applicant Name" field in the source data is mapped to the standard "Applicant Name" field, and the "Contact Mobile Phone" field is mapped to the standard "Contact Telephone" field.
[0078] Based on the established mapping and renaming relationships, entity type conversion and field name standardization are performed on each source data to obtain semantically aligned standardized data. This process involves operations such as field renaming, data format conversion, and entity type determination.
[0079] In practical applications, such as when a municipal government service center processes building permit applications, applicant information differs across systems: the housing and construction department's system refers to the applicant as "construction unit," while the planning department's system refers to the applicant as "applicant enterprise," and the field definitions and data formats are also inconsistent. The method described above can map this information from different sources to a unified "legal entity applicant" standard entity, achieving standardization of field names and data formats, and providing a foundation for subsequent cross-departmental business collaboration.
[0080] In one optional implementation, a traceability chain is established. By calculating the consistency frequency and node verification pass rate of each attribute value in the standardized data, a dynamic trustworthy metric value is assigned to each attribute value, resulting in an enhanced data representation with traceability tags and trustworthy metric values, including:
[0081] Extract each attribute value and historical correction record from the standardized data and organize them into the traceability chain in chronological order;
[0082] For each attribute value in the standardized data, retrieve historical attribute values that are the same as the current attribute value from the set of attribute values of completed approval items, and calculate the consistency frequency between the approval results corresponding to the historical attribute values and the expected approval results; for each transmission path node in the traceability chain, extract the verification records of the current transmission path node performing verification operations on the data, and calculate the node verification pass rate by the ratio between the number of successful verifications in the verification records and the total number of verifications.
[0083] The consistency frequency and the path verification pass rate are weighted and summed, and the weighted summation result is attenuated according to the number of corrections in the historical correction records to obtain a dynamic trust metric value. The dynamic trust metric value is associated with the corresponding traceability chain to the corresponding attribute value in the standardized data to obtain an enhanced data representation with traceability tags and trust metric values.
[0084] Extract each attribute value and related historical correction records from the standardized data. A typical business data record contains multiple attribute fields, such as name, contact information, and address in customer information. For each attribute value, collect all its modification history from creation to the current state, including modification time, modifier, previous value, and current value, forming a complete modification trajectory. For example, for a customer address information record, its traceability chain records the modification process from "Haidian District, Beijing" to "Chaoyang District, Beijing," including modification time and operator information.
[0085] Construct a historical attribute value database to record all attribute values that have undergone approval processes and their approval results. When a new attribute value needs to be evaluated, retrieve historical records that are the same as or similar to the current value from the database, and calculate the proportion of approval results in these historical records that meet expectations as the consistency frequency. For example, if the address value "Chaoyang District, Beijing" has been approved 95 out of 100 times in the past, its consistency frequency is 0.95.
[0086] For each transmission path node in the traceability chain, the verification records of the data verification operations performed by the current transmission path node are extracted. The ratio between the number of successful verifications in the verification records and the total number of verifications is calculated to obtain the node's verification pass rate. Transmission path nodes can be various processing steps in the data flow process, such as data entry, cleaning, and conversion. For example, if the data cleaning node performs 50 verification operations on the address field, and 48 of them are successful, then the node's verification pass rate for the address field is 0.96.
[0087] After obtaining the consistency frequency and node verification pass rate, the two are weighted and summed. The weighting coefficient can be adjusted according to the specific business scenario; for example, the weight of the consistency frequency can be set to 0.6, and the weight of the node verification pass rate to 0.4. Through weighted calculation, a preliminary credibility score is obtained.
[0088] The weighted summation result is attenuated based on the number of corrections recorded in the historical correction records. The more corrections, the lower the stability of the data, so the credibility needs to be appropriately reduced. The attenuation function can adopt an exponential decay form, for example: adjustment coefficient = exp(-0.1 × number of corrections). The initial score is multiplied by the adjustment coefficient to obtain the final dynamic credibility quantification value.
[0089] The calculated dynamic credibility metrics are associated with the corresponding traceability chain and the corresponding attribute values in the standardized data to form an enhanced data representation with traceability tags and credibility metrics. This representation not only includes the original data information, but also adds the data source, flow process and credibility assessment results, providing a more comprehensive decision-making basis for subsequent data applications.
[0090] In practical applications, when evaluating a government loan application, the information provided by the applicant can be assessed for credibility using the methods described above. This involves establishing a traceability chain for income information, recording the entire process from data submission and verification to final storage, retrieving approval information for similar cases in the past, calculating the consistency frequency, assessing the reliability of information verification at each node, and comprehensively calculating a quantifiable value for the credibility of the information. If the credibility is lower than a preset threshold, additional verification is prompted.
[0091] In this approach, data is no longer a simple static record, but an enhanced representation with comprehensive traceability information and credibility assessment results, providing a more reliable foundation for data-driven decision-making while improving the transparency and accountability of data use.
[0092] In one optional implementation, comparing attribute values of the same approval item on the enhanced data representation to identify attribute value conflicts between different source data includes:
[0093] For the same approval item, an attribute value set of the same attribute from different source data is extracted from the enhanced data representation. Each attribute value in the attribute value set is compared pairwise. When there is a numerical difference or semantic difference, it is marked as a candidate conflicting attribute value pair.
[0094] Extract inter-attribute dependency rules related to the candidate conflicting attribute value pairs from a predefined attribute semantic dependency rule base. Extract the attribute values of the associated attributes from the enhanced data representation according to the inter-attribute dependency rules. Determine whether the attribute values in the candidate conflicting attribute value pairs and the attribute values of the associated attributes satisfy the value constraint relationship. If the value constraint relationship is not satisfied, confirm the candidate conflicting attribute value pairs as real conflicting attribute value pairs.
[0095] For each attribute value in the real conflict attribute value pair, the corresponding traceability chain is extracted, and the traceability integrity of the data source identifier and the continuity integrity of the transmission path node sequence in the traceability chain are calculated. The traceability integrity and continuity integrity are weighted and summed to obtain the traceability chain integrity score. The traceability chain integrity score and the credible quantification value of the corresponding attribute value are associated with the real conflict attribute value pair to obtain the attribute value conflict.
[0096] For the same approval item, an enhanced data representation extracts a set of attribute values for the same attribute from different source data. Enhanced data representation refers to a structured data representation formed after data integration and standardization, which includes attribute names, attribute values, and traceability information. For example, for a business registration approval item, relevant data was obtained from the industrial and commercial administration department, tax department, and banking system, forming an enhanced data representation. The attribute "registered capital" is extracted from this representation, resulting in a set of attribute values: {1 million yuan, 1,000,000 yuan, 1 million}.
[0097] The system compares each attribute value in the attribute value set pairwise. When numerical or semantic differences exist, the pairs are marked as candidate conflicting attribute value pairs. Numerical differences refer to inconsistencies in the numerical values themselves, such as "100" and "200". Semantic differences refer to different expressions that refer to the same thing; semantic analysis is needed to determine if there is a substantial difference, such as "1 million yuan" and "1,000,000 yuan". The comparison methods include: for numerical attributes, unit standardization and numerical normalization are performed before directly comparing the numerical values; for textual attributes, string similarity calculations are used, such as edit distance and Jaccard similarity; for date attributes, conversion to a standard time format is performed before comparison; for enumerated attributes, it is determined whether they are different expressions of the same enumerated value.
[0098] In the above example, "1 million yuan" and "1,000,000 yuan" are confirmed to be semantically consistent after numerical normalization and do not constitute a candidate conflict; however, if there is a case of "1 million yuan" and "2 million yuan", it will be marked as a candidate conflict attribute value pair.
[0099] The system extracts attribute dependency rules related to candidate conflicting attribute value pairs from a predefined attribute semantic dependency rule base. These rules describe the value constraints between different attributes, such as the relationship between "registered capital" and "paid-in capital" (paid-in capital should not exceed registered capital). The rule base is built based on domain knowledge and includes the following rule types: numerical constraint rules (e.g., the value of attribute A should be greater than the value of attribute B); temporal constraint rules (e.g., the time of attribute A should be earlier than the time of attribute B); enumerated value constraint rules (e.g., when attribute A takes a certain value, attribute B can only take a value from a specific set of values); and conditional constraint rules (e.g., when attribute A meets a specific condition, attribute B should meet the corresponding constraint).
[0100] Based on the extracted attribute dependency rules, the attribute values of related attributes are extracted from the augmented data representation. It is then determined whether the attribute values in the candidate conflicting attribute value pairs satisfy the value constraint relationship with the attribute values of the related attributes. If the value constraint relationship is not satisfied, the candidate conflicting attribute value pairs are confirmed as actual conflicting attribute value pairs.
[0101] Taking the relationship between registered capital and paid-in capital as an example, if the registered capital extracted from different source data is "1 million yuan" and "2 million yuan" respectively (marked as candidate conflict), and the paid-in capital is "1.5 million yuan", then according to the rule that "paid-in capital should not exceed registered capital", the paid-in capital of "1 million yuan" and "1.5 million yuan" does not meet the constraint relationship. Therefore, "1 million yuan" and "2 million yuan" are confirmed as a real conflict attribute value pair.
[0102] For each attribute value in a real conflict attribute value pair, the corresponding source chain is extracted. The source chain records the complete path of data from its original generation to its current representation, including information such as data source identifier, intermediate processing nodes, and processing timestamps. The data source identifier includes system number, database name, department code, etc.; the transmission path nodes include data collection points, intermediate processing systems, data transmission gateways, etc.
[0103] The traceability integrity of the data source identifier and the continuity integrity of the transmission path node sequence are calculated in the traceability chain. The traceability integrity reflects the clarity of the data source and is calculated by dividing the number of valid data source identifiers by the total number of expected identifiers. The continuity integrity measures the degree of complete recording of the data flow process and is calculated by dividing the number of valid connections between adjacent nodes by the total number of nodes minus one.
[0104] The traceability completeness score is obtained by weighted summation of traceability completeness and continuity completeness. The weighting coefficient can be set according to the application scenario, typically with a traceability weight of 0.6 and a continuity weight of 0.4. The completeness score ranges from 0 to 1, with a higher score indicating a more complete traceability chain and higher reliability of the attribute value.
[0105] By associating the traceability chain integrity score with the corresponding credible quantification value of the attribute value to the actual conflicting attribute value pair, the final representation of the attribute value conflict is obtained. The credible quantification value is an indicator calculated based on the traceability chain integrity score and other factors (such as data source reliability rating, data timeliness, etc.).
[0106] The above methods can effectively identify attribute value conflicts in multi-source government data and provide a credibility assessment of conflicting attribute values, laying the foundation for subsequent conflict resolution and data quality improvement. In practical applications, this can be used in multi-source data fusion scenarios such as government data integration, significantly improving data consistency and reliability.
[0107] In one optional implementation, multiple attribute values are fused based on their respective credible quantification values, the integrity of the tracing chain, and the semantic dependencies between attributes to obtain a conflict-resolved fused data entity, including:
[0108] The credible quantification value associated with each attribute value in the real conflict attribute value pair and the traceability chain integrity score are extracted and weighted to obtain a comprehensive credibility score. The attribute values in the real conflict attribute value pair are sorted according to the comprehensive credibility score, and the attribute value with the highest comprehensive credibility score is selected as the initial fusion attribute value.
[0109] According to the attribute dependency rules, the attribute values of the corresponding associated attributes are extracted from the enhanced data representation. The degree of conformity between the initial fusion attribute value and the attribute values of each associated attribute is calculated to satisfy the attribute dependency rules. The degree of conformity corresponding to all attribute dependency rules is weighted and aggregated to obtain the semantic consistency score.
[0110] The overall credibility score of the initial fusion attribute value is adjusted based on the semantic consistency score and the semantic dependency strength defined by the attribute dependency rules to obtain the final fusion attribute value;
[0111] The final fused attribute value replaces the corresponding real conflict attribute value pair in the enhanced data representation, and the traceability chain and trusted quantification associated with the final fused attribute value are retained in the replaced enhanced data representation to obtain the fused data entity after conflict resolution.
[0112] For identified pairs of conflicting attribute values, the credible quantifiable values associated with each attribute value and the traceability chain integrity score are extracted and then weighted and fused. The credible quantifiable value represents the reliability of the data source, typically determined by factors such as the authority of the data source and its historical accuracy. The traceability chain integrity score reflects the transparency and traceability of the data flow process. The weighted fusion process can employ a linear weighting method, i.e.: Overall Credibility Score = α × Credible Quantifiable Value + β × Traceability Chain Integrity Score. Here, α and β are weighting coefficients, and α + β = 1. The weighting coefficients can be adjusted according to specific application scenarios to reflect the importance of different factors. For example, in government data fusion, where data traceability is more important, the β value can be appropriately increased.
[0113] After calculating the overall credibility score, the attribute values in the real conflict attribute value pairs are sorted in descending order, and the attribute value with the highest overall credibility score is selected as the initial fusion attribute value.
[0114] Considering the semantic dependencies between attributes, we extract the associated attribute values that have semantic dependencies with the initially selected fusion attribute values from the enhanced data representation. Semantic dependencies are usually predefined in the dependency rules between attributes. For example, there is a time calculation relationship between "date of birth" and "age", and a geographical inclusion relationship between "city" and "province".
[0115] For each dependency rule between attributes, the degree of conformity between the initially selected fused attribute value and the associated attribute value is calculated. The method for calculating the degree of conformity varies depending on the type of dependency rule. For example, for numerical dependencies, the degree of mathematical relationship can be calculated; for categorical dependencies, the degree of category matching can be calculated; and for textual dependencies, semantic similarity can be calculated.
[0116] The semantic consistency score is obtained by weighted aggregation of the conformity scores of all dependent rules: Semantic consistency score = Σ(wi × conformity score of rule i), where wi is the weight of the i-th rule, and Σwi = 1. The weights can be allocated according to the importance and reliability of the rules.
[0117] Based on the semantic consistency score and the strength of dependencies between attributes, the overall credibility score of the initially selected fused attribute values is adjusted: Adjusted overall credibility score = Overall credibility score × (1 + γ × Semantic consistency score × Dependency strength), where γ is a semantic adjustment factor that controls the magnitude of the impact of semantic consistency on credibility. Dependency strength reflects the strictness of semantic rules, and its value range is usually [0, 1], with a larger value indicating a stronger dependency.
[0118] After the adjustment is completed, the final fusion attribute value is obtained. The final fusion attribute value replaces the corresponding real conflict attribute value pair in the enhanced data representation, and the traceability chain and trusted quantification value associated with the fusion attribute value are retained, thus obtaining the fusion data entity after conflict resolution.
[0119] Maintaining the traceability chain ensures the transparency and traceability of the data fusion process, facilitating subsequent verification and auditing of the fusion results. Simultaneously, retaining reliable quantifiable values provides data users with a reference for assessing the quality of the fusion results.
[0120] The above method not only considers the reliability and traceability of the data source, but also incorporates semantic constraints between attributes, making the conflict resolution results more accurate and reasonable, and improving the quality and reliability of data fusion.
[0121] In one optional implementation, path reasoning is performed on the government approval knowledge graph based on graph reasoning rules to identify incomplete application materials and risk associations of related entities, resulting in a set of reasoning results including:
[0122] Extract multi-level material dependency reasoning rules and subject association propagation reasoning rules from a predefined graph reasoning rule base, and perform multi-path parallel traversal in the government approval knowledge graph starting from the approval item identification node;
[0123] Calculate the structural similarity between each traversal path and each path pattern in the multi-level material dependency reasoning rules, select the material necessity weight corresponding to the path pattern with the highest structural similarity as the weight value of the traversal path, aggregate the weight values of multiple traversal paths that reach the same material identifier node to obtain the aggregate necessity score of the material identifier node, and mark the material identifier node as a material missing item when the aggregate necessity score is lower than the preset material missing judgment threshold;
[0124] Extract approval entity identifier nodes with risk identification attributes from the government approval knowledge graph as risk source nodes. Perform breadth-first traversal along the business relationship edges starting from the risk source nodes. During the traversal, calculate the risk propagation score of the current node according to the entity association propagation reasoning rules. Extract approval entity identifier nodes with risk propagation scores higher than the preset risk association judgment threshold as risk-related entities and record the risk propagation path.
[0125] The missing material items and their corresponding aggregation necessity scores, the risk-related entities, and the risk propagation paths are encapsulated into a set of inference results.
[0126] like Figure 2 As shown, the method includes:
[0127] When performing path reasoning on the government approval knowledge graph based on graph reasoning rules, multi-level material dependency reasoning rules and subject association propagation reasoning rules are extracted from a predefined graph reasoning rule library. In the government approval knowledge graph, a multi-path parallel traversal is performed starting from the approval item identifier node. The multi-level material dependency reasoning rules include a series of path patterns, each defining a specific dependency relationship between a material node and an approval item node, as well as the corresponding material necessity weight. For example, for directly dependent materials, the path pattern can be represented as "Approval Item - Direct Dependency → Application Materials," with a corresponding weight of 0.8; for indirectly dependent materials, the path pattern can be represented as "Approval Item - Related Items → Related Items - Direct Dependency → Application Materials," with a corresponding weight of 0.6.
[0128] When performing multi-path parallel traversal, starting from the approval item identifier node, a depth-first search is performed along multiple relational edges. The node sequence and relation type of each traversal path are recorded. For each traversal path, its structural similarity with each path pattern in the multi-level material dependency inference rules is calculated. The structural similarity calculation considers node type matching degree and relation type matching degree. For paths with complete node type matching and complete relation type matching, the similarity is 1; for partially matching paths, a similarity value between 0 and 1 is assigned based on the matching degree.
[0129] The material necessity weight corresponding to the path pattern with the highest structural similarity is selected as the weight value of the traversal path. When there are multiple traversal paths leading to the same material identifier node, the weight values of these paths are aggregated to obtain the aggregate necessity score of the material identifier node. The aggregation method can be maximum value aggregation or weighted average aggregation. For example, when the weight of path one is 0.8 and the weight of path two is 0.6, the aggregate necessity score obtained by using maximum value aggregation is 0.8.
[0130] Set a preset threshold for determining missing materials, such as 0.7. When the aggregation necessity score of a material identifier node is lower than this threshold, the material identifier node is marked as a missing material item. For example, if the aggregation necessity score of a material node is 0.65, which is lower than the threshold of 0.7, it will be marked as a missing item, indicating that the material is required in the current approval item but has not been provided.
[0131] In the government approval knowledge graph, nodes with risk-identifying attributes are extracted as risk source nodes. These risk-identifying attributes may include "historical violation records" and "low credit rating." Starting from the risk source node, a breadth-first traversal is performed along the edges of business relationships, which may include "project undertaking units," "qualification affiliation units," and "joint application units."
[0132] During the traversal, the risk propagation score of the current node is calculated according to the subject association propagation inference rule. This rule defines the attenuation coefficient of risk across different relationship types. For example, for the "project undertaking unit" relationship, the attenuation coefficient can be set to 0.9; for the "qualification affiliation unit" relationship, the attenuation coefficient can be set to 0.7. The risk propagation score is calculated by multiplying the risk score of the source node by the cumulative value of the attenuation coefficients of each relationship along the path. For example, if the risk score of the source node is 0.8, and it reaches a certain node through the edges "project undertaking unit" and "qualification affiliation unit," then the risk propagation score of that node is 0.8 multiplied by 0.9 multiplied by 0.7, which equals 0.504.
[0133] A preset risk association determination threshold, such as 0.5, is set. Approval entity identifier nodes with risk propagation scores higher than this threshold are extracted as risk-associated entities, and the risk propagation path from the risk source node to the risk-associated entity is recorded, including intermediate nodes and relationship types. For example, if a certain entity node has a risk propagation score of 0.6, which is higher than the threshold of 0.5, it is marked as a risk-associated entity.
[0134] The missing materials, their corresponding aggregate necessity scores, risk-related entities, and risk propagation paths are encapsulated into a set of inference results, which serves as the basis for subsequent approval decisions. This set of inference results is organized in structured data format and includes two parts: material completeness assessment results and risk association assessment results. The material completeness assessment results include a list of missing materials and the necessity score for each material; the risk association assessment results include a list of risk-related entities and their corresponding risk propagation paths and risk propagation scores.
[0135] The above methods can effectively identify missing materials and potential risks in the government approval process, improve approval efficiency and risk control capabilities, and provide intelligent support for government services.
[0136] In one optional implementation, based on the set of inference results, combined with approval business rules and quantitative assessment criteria, the risk level of the application is quantified and a decision path is recommended to generate auxiliary decision-making results, including:
[0137] Extract the aggregation necessity score of missing materials and the risk propagation score of the risk-associated subject from the set of inference results; obtain the material type importance coefficient and calculate the material integrity risk score by weighting it with the aggregation necessity score; obtain the subject risk type weight coefficient and calculate the subject association risk score by weighting it with the risk propagation score.
[0138] The material integrity risk score and the subject association risk score are fused from multiple dimensions to obtain a comprehensive risk assessment score. The comprehensive risk assessment score is then mapped to the corresponding risk level identifier according to a preset risk level classification rule.
[0139] Based on the risk level identifier, the missing materials are prioritized according to the aggregated necessity score and a supplementary material list is generated. The risk propagation path in the inference result set is decomposed, and the business relationship edge type sequence is extracted. Based on the business relationship edge type sequence, a targeted risk verification process suggestion is generated. Based on the risk level identifier, a predefined approval process processing strategy is matched to generate an approval process optimization scheme.
[0140] The comprehensive risk assessment score, the risk level identifier, the list of supplementary materials, the risk verification process suggestions, and the approval process optimization plan are packaged into the auxiliary decision-making results.
[0141] Key evaluation indicator data are extracted from the inference result set, which typically includes application material analysis results and information on risk-related entities. The extraction process primarily targets two types of indicators: the aggregate necessity score of missing materials and the risk propagation score of risk-related entities. The aggregate necessity score reflects the impact of missing materials on the overall application evaluation and is usually pre-calculated based on business rules and historical data and stored in a knowledge graph. For example, in the construction project planning permit scenario, the aggregate necessity score of the land use planning permit is 0.85 (high necessity), while the surrounding environmental impact statement is only 0.35 (low necessity). The risk propagation score represents the transmissibility of potential risks from related entities to the applicant and is calculated using a risk propagation path algorithm.
[0142] Obtain a pre-defined material type importance coefficient table. This table defines the relative importance of various materials based on different business scenarios. For example, in business establishment permit approval, the importance coefficient for a business license is 0.9, for a business premises certificate it is 0.8, and for a supplementary commitment letter it is 0.4. Weight the aggregate necessity score of each missing material item with its corresponding material type importance coefficient to obtain the material integrity risk score. The calculation method is to multiply the aggregate necessity score of each missing material by its importance coefficient and then sum them to form the total risk score.
[0143] The system configuration retrieves a table of risk type weighting coefficients for each entity. This table defines the weights for different types of associated entity risks, such as administrative penalty risk (0.8), qualification violation risk (0.7), and credit default risk (0.6). The risk propagation score of each associated entity is multiplied by its corresponding risk type weighting coefficient to obtain a weighted associated entity risk score. When multiple associated entities exist, a weighted summation method is used to calculate the overall associated risk score.
[0144] A multi-dimensional fusion algorithm is used to calculate the comprehensive risk assessment score, integrating the material integrity risk score and the entity-related risk score. The fusion process considers the correlation and complementarity of the two risks, and weights the calculation by setting fusion weight coefficients (typically 0.4 for material integrity and 0.6 for entity-related risk). In specific business scenarios, the fusion weights can be dynamically adjusted based on historical case analysis. The comprehensive risk assessment score obtained after fusion typically ranges from 0 to 1, with higher scores indicating higher risk levels.
[0145] Based on the preset risk level classification rules, the comprehensive risk assessment score is mapped to the corresponding risk level label. Typical classification rules are: low risk (0-0.3), medium risk (0.3-0.6), high risk (0.6-0.8), and very high risk (0.8-1.0). The current application is then labeled with a risk level label, such as "medium risk" or "high risk".
[0146] Based on the identified risk level, a list of supplementary materials is generated. Missing materials are sorted in descending order of their aggregate necessity score, with priority thresholds set (e.g., high priority > 0.7, medium priority > 0.4, low priority <= 0.4). A structured list of supplementary materials is then generated based on the sorting results. The list includes information such as material name, necessity level, and submission deadline, providing applicants with clear guidance on supplementing their materials.
[0147] The risk propagation paths in the inference result set are decomposed, and the business relationship edge type sequences between each node in the path are extracted, such as "project undertaking unit → qualification affiliation unit → historical administrative penalty record". Based on the feature patterns of the edge type sequences, predefined risk verification strategy templates are matched to generate targeted risk verification process suggestions. For example, for risk paths involving qualification affiliation, it is recommended to implement a special qualification authenticity verification procedure; for risk paths involving historical penalties, it is recommended to implement a compliance rectification review process.
[0148] Based on risk level identifiers, predefined approval process strategies are matched. Different risk levels correspond to different approval levels, approval steps, and approval time limits. For example, low-risk applications can adopt simplified processes and reduce approval steps, while high-risk applications require additional steps such as expert review and cross-verification. The system-generated approval process optimization plan includes flowcharts, key node descriptions, and processing suggestions, providing decision-making references for approvers.
[0149] The comprehensive risk assessment score, risk level identification, supplementary material list, risk verification process suggestions, and approval process optimization plan are packaged into a structured auxiliary decision-making result. This result is then output to the business system or decision support platform through a standard interface for approval personnel to refer to. The decision result adopts a hierarchical structure, which makes it easy for approval personnel to view information details from different dimensions as needed.
[0150] The above methods enable risk quantification assessment and decision path recommendation based on reasoning results, providing objective and comprehensive risk identification and handling solutions for government approval processes.
[0151] This invention relates to an intelligent auxiliary decision-making system for government approval based on multi-source data fusion, the system comprising:
[0152] The first unit is used to perform structured parsing of the source data in the multi-source heterogeneous data set of government approval business, and to map the same type of entity objects in different source data to a unified semantic space to obtain standardized data after semantic alignment.
[0153] The second unit is used to establish a traceability chain. By calculating the consistency frequency and node verification pass rate of each attribute value in the standardized data, dynamic trust quantification values are assigned to each attribute value, resulting in an enhanced data representation with traceability tags and trust quantification values.
[0154] The third unit is used to compare the attribute values of the same approval item in the enhanced data representation, identify attribute value conflicts between different source data, and fuse multiple attribute values based on the credible quantification value of each conflicting attribute value, the integrity of the traceability chain, and the semantic dependency relationship between attributes to obtain a fused data entity after conflict resolution.
[0155] The fourth unit is used to construct a government approval knowledge graph by taking the fused data entities as nodes and the business relationships in government approval items as edges, and to perform path reasoning on the government approval knowledge graph based on graph reasoning rules to identify the incompleteness of application materials and the risk association of related entities, and to obtain a set of reasoning results.
[0156] The fifth unit is used to quantify the risk level and recommend decision-making paths for the application items based on the set of reasoning results, combined with the approval business rules and quantitative evaluation criteria, and to generate auxiliary decision-making results.
[0157] A third aspect of the present invention provides an electronic device, comprising:
[0158] processor;
[0159] Memory used to store processor-executable instructions;
[0160] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0161] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0162] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent auxiliary decision-making in government approval based on multi-source data fusion, characterized in that: include: The data from various sources in the multi-source heterogeneous data set of government approval business are structured and parsed. Similar entity objects in different source data are mapped to a unified semantic space to obtain standardized data after semantic alignment. The similar entity objects include enterprise applicants and enterprise registered capital. The similar entity objects contain multiple attribute fields, and the value of the attribute field is the attribute value. A traceability chain is established by calculating the consistency frequency and node verification pass rate of each attribute value in the standardized data. Each attribute value is a value for each attribute field of the same type of entity object. Dynamic trustworthy quantification values are assigned to each attribute value, resulting in an enhanced data representation with traceability tags and trustworthy quantification values, including: Extract each attribute value and historical correction record from the standardized data and organize them into the traceability chain in chronological order; For each attribute value in the standardized data, historical attribute values that are the same as the current attribute value are retrieved from the set of attribute values of completed approval items. The consistency frequency between the approval results corresponding to the historical attribute values and the expected approval results is calculated. The consistency frequency is the ratio of the number of times the approval results in the historical approval records meet the expectations to the total number of historical approvals. For each transmission path node in the traceability chain, the verification records of the current transmission path node performing verification operations on the data are extracted. The node verification pass rate is obtained by calculating the ratio between the number of times the verification passes in the verification records and the total number of verifications. The consistency frequency and the path verification pass rate are weighted and summed, and the weighted summation result is attenuated according to the number of corrections in the historical correction records to obtain a dynamic trust metric value. The dynamic trust metric value is associated with the corresponding traceability chain to the corresponding attribute value in the standardized data to obtain an enhanced data representation with traceability mark and trust metric value. The enhanced data representation is compared with the attribute values of the same approval item. The attribute values of the same approval item include each attribute value of the same type of entity object. Attribute value conflicts between different source data are identified. Based on the credible quantification value of each conflicting attribute value, the integrity of the traceability chain, and the semantic dependency between attributes, multiple attribute values are fused. The multiple attribute values refer to multiple conflicting attribute values, resulting in a fused data entity after conflict resolution. Using the fused data entities as nodes and the business relationships in government approval items as edges, a government approval knowledge graph is constructed. Based on graph reasoning rules, path reasoning is performed on the government approval knowledge graph to identify the incompleteness of application materials and the risk association of related entities, and a set of reasoning results is obtained. Based on the set of reasoning results, combined with approval business rules and quantitative evaluation criteria, the risk level of the application is quantified and a decision path is recommended to generate auxiliary decision-making results.
2. The method according to claim 1, characterized in that, The data from various sources in the multi-source heterogeneous dataset of government approval processes are structured and parsed. Similar entity objects from different sources are mapped to a unified semantic space, resulting in standardized data with semantic alignment, including: Semantic features are extracted from each source data in the multi-source heterogeneous data set. The statistical distribution characteristics of each field in each source data are analyzed. Each field is represented as a composite feature vector according to the statistical distribution characteristics. Entity objects of the same type in each source data are represented as entity feature matrices according to the corresponding composite feature vectors. In the standard semantic space, standard composite feature vectors are generated according to the attribute definition specifications of each standard entity type, and all standard composite feature vectors of the standard entity type are combined into a standard entity type prototype matrix. Calculate the matrix field similarity between the entity feature matrix of each source data and the standard entity type prototype matrix, and aggregate the matrix field similarity in the entity feature matrix to obtain the entity-level semantic mapping score; Based on the entity-level semantic mapping score, the mapping relationship between similar entity objects in each source data and standard entity types is determined. Based on the matrix field similarity, the renaming relationship between each source data field and standard field name is determined. Based on the mapping relationship and the renaming relationship, entity type conversion and field name standardization are performed on each source data to obtain semantically aligned standardized data.
3. The method according to claim 1, characterized in that, The enhanced data representation is used to compare attribute values for the same approval item, and the identification of attribute value conflicts between different source data includes: For the same approval item, the attribute value set of the same type of entity object from different source data is extracted from the enhanced data representation. Each attribute value in the attribute value set is compared pairwise. When there is a numerical difference or semantic difference, it is marked as a candidate conflicting attribute value pair. Extract inter-attribute dependency rules related to the candidate conflicting attribute value pairs from a predefined attribute semantic dependency rule base. Extract the attribute values of the associated attributes from the enhanced data representation according to the inter-attribute dependency rules. Determine whether the attribute values in the candidate conflicting attribute value pairs and the attribute values of the associated attributes satisfy the value constraint relationship. If the value constraint relationship is not satisfied, confirm the candidate conflicting attribute value pairs as real conflicting attribute value pairs. For each attribute value in the real conflict attribute value pair, the corresponding traceability chain is extracted, and the traceability integrity of the data source identifier and the continuity integrity of the transmission path node sequence in the traceability chain are calculated. The traceability integrity and continuity integrity are weighted and summed to obtain the traceability chain integrity score. The traceability chain integrity score and the credible quantification value of the corresponding attribute value are associated with the real conflict attribute value pair to obtain the attribute value conflict.
4. The method according to claim 3, characterized in that, Based on the credible quantification values of each conflicting attribute value, the completeness of the tracing chain, and the semantic dependencies between attributes, multiple attribute values are fused to obtain a conflict-resolved fused data entity, including: The credible quantification value associated with each attribute value in the real conflict attribute value pair and the traceability chain integrity score are extracted and weighted to obtain a comprehensive credibility score. The attribute values in the real conflict attribute value pair are sorted according to the comprehensive credibility score, and the attribute value with the highest comprehensive credibility score is selected as the initial fusion attribute value. According to the attribute dependency rules, the attribute values of the corresponding associated attributes are extracted from the enhanced data representation. The degree of conformity between the initial fusion attribute value and the attribute values of each associated attribute is calculated to satisfy the attribute dependency rules. The degree of conformity corresponding to all attribute dependency rules is weighted and aggregated to obtain the semantic consistency score. The overall credibility score of the initial fusion attribute value is adjusted based on the semantic consistency score and the semantic dependency strength defined by the attribute dependency rules to obtain the final fusion attribute value; The final fused attribute value replaces the corresponding real conflict attribute value pair in the enhanced data representation, and the traceability chain and trusted quantification associated with the final fused attribute value are retained in the replaced enhanced data representation to obtain the fused data entity after conflict resolution.
5. The method according to claim 1, characterized in that, Based on graph reasoning rules, path reasoning is performed on the government approval knowledge graph to identify incomplete or missing application materials and risk associations with related entities, resulting in a set of reasoning results including: Extract multi-level material dependency reasoning rules and subject association propagation reasoning rules from a predefined graph reasoning rule base, and perform multi-path parallel traversal in the government approval knowledge graph starting from the approval item identification node; Calculate the structural similarity between each traversal path and each path pattern in the multi-level material dependency reasoning rules, select the material necessity weight corresponding to the path pattern with the highest structural similarity as the weight value of the traversal path, aggregate the weight values of multiple traversal paths that reach the same material identifier node to obtain the aggregate necessity score of the material identifier node, and mark the material identifier node as a material missing item when the aggregate necessity score is lower than the preset material missing judgment threshold; Extract approval entity identifier nodes with risk identification attributes from the government approval knowledge graph as risk source nodes. Perform breadth-first traversal along the business relationship edges starting from the risk source nodes. During the traversal, calculate the risk propagation score of the current node according to the entity association propagation reasoning rules. Extract approval entity identifier nodes with risk propagation scores higher than the preset risk association judgment threshold as risk-related entities and record the risk propagation path. The missing material items and their corresponding aggregation necessity scores, the risk-related entities, and the risk propagation paths are encapsulated into a set of inference results.
6. The method according to claim 1, characterized in that, Based on the aforementioned set of reasoning results, and in conjunction with approval business rules and quantitative assessment criteria, the risk level of the application is quantified and a decision-making path is recommended, generating auxiliary decision-making results including: Extract the aggregation necessity score of missing materials and the risk propagation score of the risk-associated subject from the set of inference results; obtain the material type importance coefficient and calculate the material integrity risk score by weighting it with the aggregation necessity score; obtain the subject risk type weight coefficient and calculate the subject association risk score by weighting it with the risk propagation score. The material integrity risk score and the subject association risk score are fused from multiple dimensions to obtain a comprehensive risk assessment score. The comprehensive risk assessment score is then mapped to the corresponding risk level identifier according to a preset risk level classification rule. Based on the risk level identifier, the missing materials are prioritized according to the aggregated necessity score and a supplementary material list is generated. The risk propagation path in the inference result set is decomposed, and the business relationship edge type sequence is extracted. Based on the business relationship edge type sequence, a targeted risk verification process suggestion is generated. Based on the risk level identifier, a predefined approval process processing strategy is matched to generate an approval process optimization scheme. The comprehensive risk assessment score, the risk level identifier, the list of supplementary materials, the risk verification process suggestions, and the approval process optimization plan are packaged into the auxiliary decision-making results.
7. An intelligent auxiliary decision-making system for government approval based on multi-source data fusion, used to implement the method as described in any one of claims 1-6, characterized in that, include: The first unit is used to perform structured parsing of the source data in the multi-source heterogeneous data set of government approval business, and to map the same type of entity objects in different source data to a unified semantic space to obtain standardized data after semantic alignment. The same type of entity objects include enterprise applicants and enterprise registered capital. The same type of entity objects contain multiple attribute fields, and the value of the attribute field is the attribute value. The second unit is used to establish a traceability chain. By calculating the consistency frequency and node verification pass rate of each attribute value in the standardized data, where each attribute value is a value of each attribute field of the same type of entity object, dynamic trust quantification values are assigned to each attribute value to obtain an enhanced data representation with traceability tags and trust quantification values. The third unit is used to compare the attribute values of the same approval item in the enhanced data representation. The attribute values of the same approval item include each attribute value of the same type of entity object. It identifies attribute value conflicts between different source data, and fuses multiple attribute values based on the credible quantification value of each conflicting attribute value, the integrity of the traceability chain, and the semantic dependency relationship between attributes. The multiple attribute values refer to multiple conflicting attribute values, and a fused data entity after conflict resolution is obtained. The fourth unit is used to construct a government approval knowledge graph by taking the fused data entities as nodes and the business relationships in government approval items as edges, and to perform path reasoning on the government approval knowledge graph based on graph reasoning rules to identify the incompleteness of application materials and the risk association of related entities, and to obtain a set of reasoning results. The fifth unit is used to quantify the risk level and recommend decision-making paths for the application items based on the set of reasoning results, combined with the approval business rules and quantitative evaluation criteria, and to generate auxiliary decision-making results.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.