An AI Automatic Review System and Method for Government Affairs

Through the structured processing of government service application materials and the calculation of the complexity of rules application, a multi-application correlation map is built, and the problems of insufficient data inconsistency and rule matching accuracy in the AI ​​automatic audit system of government affairs are solved, and more efficient risk analysis and audit management are achieved.

CN120031020BActive Publication Date: 2025-07-04SHENZHEN ZHONGJING ZHENGTONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510510439.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-04
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In the existing AI automatic review system for government affairs, there is inconsistency in the internal processing of application data, insufficient accuracy in matching rules, and failure to effectively identify the application situation of complex rules, resulting in errors in terms of adaptation or missed judgments.

Method used

By identifying and structuring key elements in government service application materials, establishing standardized application data units, calculating the complexity factor of the rules, building a multi-application correlation map, analyzing the density of cross-application correlations, and generating a comprehensive audit risk warning.

Benefits of technology

It improves the consistency and completeness of application data, improves the accuracy of rule matching, enhances the depth and relevance of cross-application risk analysis, improves the pertinence and accuracy of risk warnings, and strengthens the intelligent management capabilities of government affairs review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031020B_ABST
    Figure CN120031020B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of public services, and specifically to a government affairs business AI automatic review system and method. The system includes an application information structuring module, which, based on government affairs service application materials, identifies and extracts applicant types, business types, regions, amounts, contact information, addresses, and guarantor elements in forms and text fields. In the present invention, by structurally identifying each application element such as business type, amount, contact information, address, applicant type, region, and guarantor in government affairs service application materials one by one, and mapping the elements to a unified data field, a standardized application data unit is formed, improving the consistency and integrity of application data. At the same time, in the database clause retrieval link, the rule application complexity factor is calculated based on the number of matching clauses and the nested hierarchy between the application data and the rule clauses, significantly improving the accuracy of the application and clause matching process and enabling more effective identification of complex rule application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of public service technology, and in particular to an AI automatic review system and method for government affairs. Background Art

[0002] AI automatic review of government affairs is an intelligent government service review system that uses artificial intelligence and data processing technology to realize automatic identification of government application materials, structured data conversion, rule matching verification, correlation analysis between applications and automatic risk assessment. Its main purpose is to efficiently handle government affairs, improve review speed and accuracy, and reduce the intensity of manual intervention.

[0003] In the existing technology, only the surface fields of the materials are preliminarily classified and identified, without in-depth structural transformation and comprehensive matching of standard fields, which leads to inconsistency in the internal processing of application data, and then the subsequent rule matching stage is inaccurate; in addition, when matching rule clauses, the existing technology does not effectively consider the complexity of the nested relationship between clauses, resulting in insufficient judgment of the application scenarios of complex rules, which often leads to clause adaptation errors or missed judgments. Therefore, improvement is needed. Summary of the invention

[0004] The purpose of the present invention is to solve the shortcomings existing in the prior art and to propose an AI automatic review system and method for government affairs.

[0005] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: A government affairs AI automatic review system comprises:

[0006] The application information structuring module, based on the government service application materials, identifies and extracts the applicant type, business type, region, amount, contact information, address and guarantor elements in the form and text fields, maps the elements to the preset data fields and establishes internal association identifiers to establish standardized application data units;

[0007] The rule compliance verification module searches the database for rule condition clauses matching the application business type and qualification based on the standardized application data unit, obtains the rule application complexity factor based on the number of matching clauses and the nesting level, compares the application data with the clause constraint value, time limit and geographical scope item by item based on the rule application complexity factor, determines the attribute compliance status, and generates a single application compliance status determination;

[0008] The associated network construction module, based on multiple standardized application data units, searches for applicant, associated address, and contact entity shared among different government service applications, calculates the cross-application association tightness index according to the number and category of shared entities, and constructs a graph structure with application entities as nodes and sharing relationships as edges based on the cross-application association tightness index to establish a multi-application association map;

[0009] The abnormal pattern recognition and review suggestion module, based on the multi-application association map, combines the compliance status determination of individual applications in associated applications to calculate a risk score and generate a comprehensive review risk prompt.

[0010] Preferably, the steps for obtaining the standardized application data unit are as follows: Based on government service application materials, scan and identify the content of form fields in the materials item by item, perform field category recognition and matching, and perform text field content splitting, and extract applicant type, business type, region, amount, contact information, address, and guarantor elements one by one to generate the original element set of government service application materials;

[0011] Based on the original element set of government service application materials, perform mapping matching verification on each field one by one, and perform structure format conversion and field reconstruction on the field content with successful matching to generate a set of standardized mapping fields for government service application materials;

[0012] Based on the set of standardized mapping fields for government service application materials, perform internal association matching between the standardized mapping fields, establish internal association relationship identifiers according to the matching results, and perform index construction and relationship binding of the standardized mapping fields to generate standardized application data units.

[0013] Preferably, the steps for obtaining the rule application complexity factor are as follows: Based on the application business type field and application qualification field in the standardized application data unit, extract the original field values and match them with the field definition library for field content normalization, complete the missing fields and unify the field format and semantic expression to obtain a combined set of normalized application business type fields and application qualification fields;

[0014] According to the combined set of normalized application business type fields and application qualification fields, sequentially retrieve the rule condition clauses in the database that match the business type, and screen the clauses that match the qualification fields and the combined set, and extract the nested levels, logical judgment quantities, clause reference times, clause activation period coverage, clause application area overlap degree, and call frequency of each clause to generate a set of structured information for rule condition clause matching;

[0015] Based on the set of structured information for rule condition clause matching, calculate the rule application complexity factor.

[0016] Preferably, the steps for obtaining the compliance status determination of a single application are as follows: Based on the rule application complexity factor, the application business type field value and the clause constraint value in the standardized application data unit are called item by item to perform conditional threshold comparison, field numerical range check, and timeliness matching judgment between the field value and the constraint value, forming an intermediate result set of the comparison between the application data and the clause constraints;

[0017] According to the intermediate result set of the comparison between the application data and the clause constraints, the geographical area field value in the standardized application data unit is obtained, and item-by-item mapping of the application geographical area field and the applicable geographical range of the rule clauses is performed to screen the clauses with successful geographical mapping matches, generating a matching result set of the application data and the geographical range;

[0018] Based on the matching result set of the application data and the geographical range, the compliance status of each item of application data is determined, and the compliance status determination of a single application is generated.

[0019] Preferably, the steps for obtaining the cross-application association tightness index are as follows: Based on multiple standardized application data units, the applicant field, the associated address field, and the contact information field in each standardized application data unit are extracted, and entity classification matching and deduplication processing are performed on all application data according to the field type, generating an application shared entity set;

[0020] According to the application shared entity set, the number of occurrences of the entity in different government service applications is counted item by item and the entity category to which it belongs is marked, and the repeated distribution information of each type of shared entity in each standardized application data unit is recorded, obtaining a shared entity distribution structure set;

[0021] Based on the shared entity distribution structure set, the cross-application association tightness index is calculated.

[0022] Preferably, the steps for obtaining the multi-application association graph are as follows: Based on the cross-application association tightness index, the applicant field, the associated address field, and the contact information field in multiple standardized application data units are extracted, and entity normalization identification coding and entity type annotation are performed on the content of each field, generating an application entity node set;

[0023] According to the application entity node set, the strength value of the shared relationship between application entities is calculated;

[0024] Based on the strength value of the shared relationship, a graph structure with application entities as nodes and the strength value of the shared relationship as edge weights is constructed, and an entity edge connection relationship mapping table between nodes is established, generating a multi-application association graph.

[0025] Preferably, the step of obtaining the comprehensive review risk prompt is as follows: Based on the multi-application association graph, extract the bidirectional paths of all application entity nodes from the graph structure, and retrieve the single-application compliance status determination results, path lengths, out-degrees of path end nodes, the number of fields of path start nodes, and the number of intermediate nodes passed by each path, and generate a multi-application path attribute set;

[0026] According to the multi-application path attribute set, calculate the risk scores of the corresponding entity nodes;

[0027] Based on the risk scores, divide all application entity nodes into risk level intervals according to the risk scores, and mark the risk level color coding and risk warning node numbers in each application path to generate a comprehensive review risk prompt.

[0028] The present invention provides a government service AI automatic review method, including the following steps:

[0029] Based on the government service application materials, identify and extract the applicant type, business type, region, amount, contact information, address, and guarantor elements in the form and text fields, attribute the elements to the preset data fields through field mapping, and establish the association identifiers between the elements to generate a standardized application data unit;

[0030] Based on the standardized application data unit, retrieve the database, match the rule condition clauses of the application's business type and qualifications, calculate according to the number of clause matches and nested levels to obtain the rule application complexity factor, and based on the rule application complexity factor, compare each application data with the clause constraint values, time limit, and regional scope to determine whether each attribute meets the requirements and generate the compliance status determination of a single application;

[0031] Based on multiple standardized application data units, find the applicants, addresses, and contact information entities shared across applications, calculate the relationship strength between the number of shared entities and categories to obtain the cross-application association tightness index, and based on the cross-application association tightness index, construct a graph structure with application entities as nodes and shared relationships as edges to generate a multi-application association graph;

[0032] Based on the multi-application association graph, combine the compliance status of a single application for each application, calculate the risk scores, evaluate the risks across applications, and generate a review risk prompt;

[0033] Analyze the comprehensive review risk prompt, refine the review opinions, and provide guidance for manual review and automatic decision-making to obtain review suggestions.

[0034] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0035] In the present invention, by structurally identifying each application element in the application materials for government services, such as business type, amount, contact information, address, applicant type, region, and guarantor, and mapping the elements to unified data fields, a standardized application data unit is formed, improving the consistency and integrity of the application data. At the same time, in the database clause retrieval link, the applicable complexity factor of the rule is calculated based on the number of matching clauses and the nested level between the application data and the rule clauses, significantly improving the accuracy of the application and clause matching process and enabling more effective identification of complex rule application scenarios. In addition, by analyzing the sharing relationships among different applications in entities such as applicants, addresses, and contact information, the cross-application association tightness index is calculated, and a graph structure with entities as nodes is constructed based on the index to quantitatively analyze potential risk association patterns, effectively enhancing the depth and relevance of risk analysis among cross-applications. Furthermore, the risk score is calculated by combining the graph structure with the compliance status determination of individual applications to generate a comprehensive audit risk prompt, improving the pertinence and accuracy of the risk prompt, thereby strengthening the risk management ability of government audits and realizing the intelligent upgrade of the government business inspection process. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a system flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0038] In this application, the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing behaviors within the scope authorized by laws and regulations and the personal information subject.

[0039] Please refer to Figure 1 , the present invention provides a technical solution: a government business AI automatic audit system includes:

[0040] An application information structuring module, based on the application materials for government services, identifies and extracts elements such as applicant type, business type, region, amount, contact information, address, and guarantor in the form and text fields, maps the elements to preset data fields and establishes internal association identifiers, and establishes a standardized application data unit;

[0041] The rule compliance verification module retrieves the rule condition clauses in the database that match the application business type and qualifications based on the standardized application data unit, calculates the rule application complexity factor according to the number of matching clauses and the nesting level, and based on the rule application complexity factor, compares the application data with the clause constraint values, time limits, and geographical scopes item by item to determine the compliance status of the attributes and generate a single application compliance status determination.

[0042] The associated network construction module searches for the applicant, associated address, and contact entity shared among different government service applications based on multiple standardized application data units, calculates the cross-application association tightness index according to the number and category of shared entities, and based on the cross-application association tightness index, constructs a graph structure with application entities as nodes and shared relationships as edges to establish a multi-application association map.

[0043] The abnormal pattern recognition and review recommendation module calculates the risk score based on the multi-application association map and the single application compliance status determination of the associated applications, and generates a comprehensive review risk prompt.

[0044] The steps for obtaining the standardized application data unit are as follows: Based on the government service application materials, scan and identify the form field content in the materials item by item, perform field category identification and matching, and perform text domain content splitting, and extract the applicant type, business type, region, amount, contact information, address, and guarantor elements one by one to generate the original element set of the government service application materials.

[0045] Based on the original element set of the government service application materials, perform mapping matching verification on the fields one by one, and perform structure format conversion and field reconstruction on the field content with successful matching to generate the standardized mapping field set of the government service application materials.

[0046] Based on the standardized mapping field set of the government service application materials, perform internal association matching among the standardized mapping fields, establish internal association relationship identifiers according to the matching results, and perform index construction and relationship binding of the standardized mapping fields to generate the standardized application data unit.

[0047] Specifically, based on the application materials for government services, first, scan and identify the existing field information in each material, and compare the literal meanings one by one in combination with a pre-organized list of field definitions. Accurately match the field names such as "applicant type", "business type", "region", "amount", "contact information", "address", "guarantor", etc. that appear in the list. When the matching degree is higher than a certain threshold, it is determined to belong to this field category. The threshold can be set to 0.8, and the value is obtained by statistically averaging the similarity and the minimum matching interval after word segmentation comparison of 200 historical government service application materials. If numerical information is found in the text field during the scanning and identification process, detect its format and number of digits. For example, for the amount field, it can be required to be within the range of 1 to 15 digits. Numerical values outside this range will be marked as abnormal data for further review during subsequent screening. Then continue to judge whether the contact information field is between 8 and 13 digits. If it exceeds this range, it will be recorded as possibly incorrect or incomplete. For the address or region field, refer to the city name, district and county name with a known list of geographical codes. The city and district and county codes in the list are obtained by retrieving the latest administrative division data. At the same time, considering that there are aliases or abbreviations for the names of some regions, when the string similarity between the alias and the official name is greater than 0.7, they will be included in the same geographical area for management. After these identifications and matches are completed, all the confirmed text and numerical information are classified and saved according to the field type. Compare the still unclear paragraphs or the content that fails to match the specific field definition again. If the corresponding category still cannot be found, it will be included in the temporary storage list. After the scanning is completed, all the confirmed and unconfirmed content will be sorted out and summarized, and finally the original element set of the government service application materials will be obtained.

[0048] Based on the original element set of government service application materials, perform mapping matching verification on all content with field category labels. Here, a reference table containing standard field names and specification formats is used for comparison item by item. The field names and format descriptions in this reference table are obtained by sorting through more than 300 government service application materials of different business types. When the coincidence degree between the element name and the standard name in the reference table is higher than 0.85, it is considered that direct mapping can be performed. For numerical type content such as amount or contact information, it is necessary to further verify whether it meets the preset numerical range. For example, the amount should be between 0 and 999999999. These upper and lower limits are determined by statistically calculating the maximum value of previous government service fees and leaving a certain safety margin. If an application amount exceeds this range, structural conversion will not be performed temporarily and it will be retained in the exception record list. For address and regional elements, corresponding conversion will be performed according to the regional codes in the national standard GB / T2260. If no matching item is found in the coding library after comparison, this field will be regarded as a non-standard address and marked with an additional identifier. After the above comparison, structural format conversion will be performed, and each element will be converted into a unified type. For example, phone numbers or mobile phone numbers will be uniformly processed as pure digital strings and the number of digits will be recorded, and the amount will be uniformly recorded in cents of RMB. At the same time, text items such as business type and applicant type will be unified in case. All fields that pass the mapping check and format conversion will be recombined into a standardized data record, and finally a standardized mapping field set of government service application materials will be generated.

[0049] Based on the standardized mapping field set of government service application materials, perform matching determination on the internal association relationships between fields. First, classify the same type of information into personal information groups such as "applicant type", "guarantor information", "contact information", etc., classify "business type", "amount", etc. into business data groups, and classify "region", "address", etc. into geographical mapping groups. If it is found that the applicant type field and the guarantor information field are exactly the same in the ID number or other identification codes, it is considered that there is a strong association. In the business data group, if it is found that the minimum fee threshold corresponding to the business type is 100 RMB, this value is set by selecting the part with the highest frequency in the low-amount part of the real government service payment data distribution in the previous year as a reference. Amounts below 100 RMB will be listed separately to distinguish their particularity. Then, compare the regional codes contained in the geographical mapping group. If the same regional code is repeatedly matched to the same personal information group, it is regarded as sharing geographical information. Integrate and arrange these matching results to generate an index structure that can be quickly retrieved, and record the association links between different fields. Finally, a standardized application data unit will be generated.

[0050] The steps for obtaining the rule application complexity factor are as follows: Based on the application business type field and the application eligibility field in the standardized application data unit, extract the original field values and perform a normalized matching of the field content with the field definition library to complete the missing fields and unify the field formats and semantic expression methods, obtaining a combined set of normalized application business type fields and application eligibility fields;

[0051] According to the combined set of normalized application business type fields and application eligibility fields, sequentially retrieve the rule condition clauses in the database that match the business type, and filter the clauses that match the eligibility field and the combined set. Extract the nested levels, the number of logical judgments, the clause reference times, the coverage range of the clause activation period, the clause application geographical overlap degree, and the call frequency of each clause to generate a structured information set of rule condition clause matches;

[0052] Based on the structured information set of rule condition clause matches, calculate the rule application complexity factor. The calculation formula is:

[0053] ;

[0054] where, is the rule application complexity factor, is the number of logical judgments of the th clause, is the nested level of the th clause, is the clause application geographical overlap degree of the th clause, is the clause reference times of the th clause, is the historical call frequency of the th clause, is the number of clauses in the structured information set of rule condition clause matches.

[0055] Specifically, based on the application business type field and the application eligibility field in the standardized application data unit, the missing fields are supplemented, and the field formats and semantic expression methods are unified. In the specific implementation process, first, select the standard string descriptions corresponding to entries such as "business type" and "application eligibility" from the previously obtained field definition library, and compare each original record of the application business type field and the application eligibility field item by item. If a field value is found to be missing in a certain record, search for the text with the highest similarity in the previously recorded comparison table for substitution, and record the substitution process in the missing field completion list. For numerical missing fields, a numerical mapping rule needs to be set additionally. For example, if a data entry should contain an amount value but is blank, refer to the statistical results of the historical data distribution of the same business type, take one-third of the difference between the minimum value and the maximum value as the base value. The determination of this one-third ratio comes from the range evaluation of the past 1000 similar business data, and select the median in this interval as the temporary filling value for subsequent precise verification. When unifying the field format, refer to the "Application Eligibility Noun Comparison Table" for the text information in the application eligibility field. The standard nouns in this table come from 200 publicly available government service regulation texts. When the similarity between the phrases appearing in the record and the standard nouns is higher than 0.8, consider them the same and replace them with the standard nouns. For the field values of the numerical type, they are retained in pure numerical form. For example, extra characters are removed from the phone number, and the amount is uniformly converted to RMB cents. If the numerical value of the amount entry exceeds the established interval of 0 to 1000000000, which is obtained based on the maximum and minimum values of the historical business application amounts and an additional 10% safety range, it is regarded as abnormal and marked. When unifying the semantic expression method, if aliases or synonymous phrases appear in the descriptions of the same business type field and the keyword token matching degree is higher than 0.85, they are combined into one category. After organizing all the above operations, a normalized combination set of the application business type field and the application eligibility field is finally obtained.

[0056] According to the combination of the standardized application business type field and the application qualification field, the rule condition clauses matching the business type in the database are retrieved in turn, and the clauses matching the qualification field and the combination set are screened. During the execution process, the titles and key tags of all archived rule texts are first retrieved in the local database with the business type as the index. The data of these titles and tags are derived from the summary of rule files released in the past five years. Each rule text contains several clause tags. After the rule text matching the current business type tag is retrieved, the clause paragraphs related to the application qualification field are selected, and these clause paragraphs are processed into sentences. Combined with the "Qualification Field Comparison List", the keywords in each sentence are matched with the standardized application qualification field. If the keyword similarity exceeds 0.75, the clause will be included in the candidate range, and the nesting level of the clause is recorded at the same time. This level is based on the clause tag. The hierarchical structure of titles and subtitles is counted. If a clause is found to contain multiple levels of nested content, the number of nested levels is automatically counted. The number of logical judgments is detected by detecting the number of keywords such as "and", "or", and "parallel requirements" that appear in the clause sentences. The number of clause references can be obtained by querying the frequency of occurrence of the clause ID in the rule file. The coverage of the clause activation period refers to the time interval formed by the announcement date and the abolition date. If a clause is only valid during the annual budget preparation period, it is marked as covering a specific period. The overlap of the applicable regions of the clause is compared between the place names listed in the clause and the geographical identifiers corresponding to the application region fields. If the place names and administrative levels are the same, the overlap is 1. If only the province or city level is the same, the overlap is recorded as 0.5. The call frequency statistics are from the call ID records of the daily access log. After the above recognition results are sorted out, a structured information set of rule condition clause matching is generated.

[0057] formula: The benefit of the formula is that the participation of multiple parameters makes the measurement of the matching complexity of the rule terms more three-dimensional, and can coordinately consider multiple indicators such as the number of logical judgments and the nesting level, and incorporate regional overlap, the number of clause citations and the historical call frequency into the calculation, which is conducive to showing the comprehensive difficulty of the rule terms in actual application.

[0058] The obtaining steps are as follows: First, the clause text is segmented into sentences, and the occurrence times of logical keywords such as "and", "or", "and", "negation" are identified. Subsequently, each occurrence of the keyword is regarded as a logical branch unit for summarization. The logical keyword list comes from a comparison table formed by sorting out common expressions in various rule documents. By counting the logical nodes corresponding to each sentence in the clause and accumulating, the initial value of the logical judgment quantity is obtained. If both "and" and "or" appear in a clause, they are combined for two-branch counting. The following gives an obtaining example: After performing sentence statistics on a certain clause text, three logical nodes are obtained, and then adding 1 to its attached restrictive conditions, four logical judgment values are obtained. Therefore 。

[0059] The obtaining steps are as follows: Read the hierarchical information in the clause structure. By parsing the hierarchical numbers of the clause title and subheadings, if the title number depth is at the first level, then take 1. If the clause still has the next-level subheadings, then accumulate 1 upward. The hierarchical depth quantity is finally accumulated to obtain the value. The following gives an obtaining example: It is retrieved that a certain rule clause is under the third-level subheading, and the hierarchical number of this clause is 3. Therefore 。

[0060] The obtaining steps are as follows: By identifying the geographical scope involved in the clause and matching it with the previously obtained application geographical field, assign 1 to the situation where the city and county are the same and the administrative levels are the same, assign 0.5 to the situation where only the city levels are the same but the counties are different, assign 0.2 to the situation where only the provincial levels are the same, and assign 0 if there is no overlap in the provincial and city levels. After extracting all the overlap degree values and performing weighted averaging on them, determine the final value. The following gives an obtaining example: When a certain clause takes into account two cities and is completely consistent with the application geographical field at the city level, and the two matching results are both 1, the average value obtained is 1. Therefore 。

[0061] The obtaining steps are as follows: By counting the total number of times this clause is called by different government affairs scenarios in the historical document reference table, the statistical method is to search for the unique identifier ID corresponding to this clause and count all the occurrence positions found. The following gives an obtaining example: If a certain clause appears 15 times in the file records in the past year, then 。

[0062] The acquisition steps are as follows: read the frequency information of each specific call to this clause in the system call log, first accumulate the number of calls to the same clause within any time period in a day to obtain the daily call frequency, and then accumulate the daily call frequencies to obtain the monthly call frequency. The following gives an acquisition example: if the average monthly call frequency of this clause is found to be 30 in the quarterly statistics, then 。

[0063] The acquisition steps are as follows: when generating the structured information set of rule condition clause matching previously, if it is found that there are multiple clauses matching this business type, then add up the number of all successfully matched clauses to obtain ; the following gives an acquisition example: when the number of successfully matched clauses is 6, then 。

[0064] Calculation process:

[0065] Set the number of clauses that have been matched , and the parameters of the first clause below it are respectively 、 、 、 、 , and the parameters of the second clause are respectively 、 、 、 、 ;

[0066] First, calculate the part in the brackets of the first clause:

[0067] ;

[0068] , , , take the absolute value and it is still 0.06248, ;

[0069] Therefore, the part in the brackets of the first clause is:

[0070] ;

[0071] Then, calculate the part in the brackets of the second clause:

[0072] ;

[0073] , , ,

[0074] Therefore, the part in the brackets of the second clause is:

[0075] ;

[0076] Next, perform a sum of squares and take the square root in the general formula:

[0077] ;

[0078] The result shows that when the clause is complex in terms of the number of logical judgments and nesting levels, and the citation times and call frequencies are high, the rule application complexity factor will have a relatively large value. When the value is 1.744, it indicates that the rule clause has a certain complexity in multi-dimensional matching and consideration in this case. If it is greater than 2.0, it reflects a higher degree of coupling, and if it is less than 1.0, it usually means that the rule clause is relatively simple. The subsequent links can identify and distinguish the difficulty range of inspection and processing based on this.

[0079] The steps to obtain the compliance status determination of a single application are as follows: Based on the rule application complexity factor, call the application business type field values and clause constraint values in the standardized application data unit item by item, perform conditional threshold comparison, field value range inspection, and timeliness matching judgment between the field values and the constraint values to form an intermediate result set of the comparison between the application data and the clause constraints;

[0080] According to the intermediate result set of the comparison between the application data and the clause constraints, obtain the regional field values in the standardized application data unit, perform item-by-item mapping of the application regional fields and the applicable regional scope of the rule clauses, screen the clauses with successful regional mapping matches, and generate a matching result set of the application data and the regional scope;

[0081] Based on the matching result set of the application data and the regional scope, determine the compliance status of each application data and generate the compliance status determination of a single application.

[0082] Specifically, based on the rule application complexity factor, first select the application business type field value and the clause constraint value in the standardized application data unit for item-by-item comparison. For the convenience of subsequent screening, each business type field value and the relevant clause constraint value are respectively summarized according to numerical type or literal type. If it is of numerical type, compare it with the pre-established range. For example, it is set that the amount is between 0 and 9,999,999, and the upper and lower limits of this range are determined by extracting the minimum and maximum values from the previous five hundred real inspection records and adding an additional ten percent safety margin. If it is found that the amount field exceeds 9,999,999 or is less than 0, it is regarded as not meeting the specified range and this phenomenon is recorded. At the same time, for the field values of logical attributes, perform threshold matching according to the conditional operators listed in the clause constraint value. For example, if the clause requires that the amount must be greater than 5,000, then as long as it is found during the comparison that the field value is not greater than 5,000, it is classified as not meeting the requirement. Here, 5,000 is obtained by rounding up the average cost statistics of ten business regions. After completing the comparison of amount categories, it is also necessary to check the time validity period. When the effective period or restricted period is defined in the clause constraint value, it is necessary to compare the application submission time with the overlap degree of this period. For example, the clause states that it is only valid within 180 days after the registration date, and the number of days of these 180 days is deduced from the inspection efficiency statistical data of the current region. If the application submission time has no intersection with the period specified in the clause, it is marked as not meeting the requirement. All these determinations are carried out on the basis of strictly following the clause content, forming a corresponding field result record for each comparison process, and finally summarizing all determinations to obtain the intermediate result set of the comparison between the application data and the clause constraints.

[0083] According to the intermediate result set of the comparison between the application data and the clause constraints, obtain the geographical region field value in the standardized application data unit and map it item-by-item to the geographical scope applicable to the rule clause. In specific implementation, first select all the records marked as pending geographical determination from the comparison intermediate results. These records contain the clauses that meet the business type or numerical requirements but have not yet confirmed geographical restrictions. Then read the geographical region field value and convert it into the corresponding city code or district / county identifier. This city code information refers to the latest national administrative division data and is mapped with the internal comparison table. If the matched city code is exactly the same as the regional scope listed in the clause, the geographical mapping determination of this clause is recorded as successful. If only the upper-level administrative region or part of the administrative region is matched, it is necessary to perform a second comparison according to the internally set overlap threshold. For example, the overlap threshold is set to 0.8, and this value is obtained by summarizing the coincidence of three hundred historical geographical names and the regions set in the clause. When the overlap degree is above 0.8, it can be included in the category of successful geographical mapping matching. If the overlap degree is below 0.8, it is classified as not meeting the requirements. After all candidate clauses have completed geographical comparison, the clauses with successful geographical mapping can be screened out, and at the same time, add a new association mark to these clauses, and summarize to form the matching result set of the application data and the geographical scope.

[0084] Based on the set of matching results of application data and geographical scope, check each item of application data in each record one by one to see if both the business type field value and the geographical field value have achieved a match. If it is recorded as not meeting the scope requirements or time limit in the previous numerical comparison, it is excluded. If the geographical mapping shows a complete correspondence with the terms, it is retained. When the secondary confirmation of all records is completed, it is obtained whether each application data has the compliance attributes under the current terms. Here, additional fields for specific business scenarios will be checked. For example, in the business of the cultural field, the terms may include requirements for professional qualification certificates. If no relevant certificate information is found in the previous comparison process, it is judged again here whether it is missing. At the same time, for the terms that do not conform to this business scenario, there is no need to perform matching verification anymore. After such exclusions one by one, integrate all the data items with compliance attributes to generate a determination of the compliance status of individual applications.

[0085] The steps to obtain the cross-application association tightness index are as follows: Based on multiple standardized application data units, extract the applicant field, associated address field, and contact information field in each standardized application data unit, and perform entity classification matching and deduplication processing on all application data according to the field type to generate a set of application sharing entities;

[0086] According to the set of application sharing entities, count the number of occurrences of entities in different government service applications item by item and mark the categories to which the entities belong, and record the repeated distribution information of each type of shared entity in each standardized application data unit to obtain a set of shared entity distribution structures;

[0087] Based on the set of shared entity distribution structures, calculate the cross-application association tightness index. The calculation formula is:

[0088] ;

[0089] Among them, is the cross-application association tightness index, is the total number of entity categories in the set of shared entity distribution structures, is the application frequency of the th type of shared entity, is the application cross quantity of the th type of shared entity, is the number of sources of government service applications of the th type of shared entity, is the number of matching failures of the th type of shared entity in different standardized application data units, is the number of field positions where the th type of shared entity co-occurs in all applications, is the number of changes in the field content of the th type of shared entity.

[0090] Specifically, based on the applicant field, associated address field, and contact information field in multiple standardized application data units, first confirm that each standardized application data unit contains these three types of fields. If there are conflicts or blanks in the field names, refer to a pre-organized field mapping description for comparison. Aggregate applicant names or ID numbers with consistent formats into the same identifier. For the associated address field, match it through street and area codes. For the contact information field, determine whether the digit length is between 8 and 13 digits. If it is not within this range, consider it incomplete information and record it in the incomplete list. After aggregation, remove duplicates from field entries with the same source. For example, when the same address appears repeatedly in different application data, if the address is the same as the standard city code stored in the system and the street name similarity is greater than 0.8, it is recognized as the same address entity. Here, 0.8 is selected through the statistical distribution of string similarities of 2000 address samples in the previous year. To identify duplicate contact information, the last four digits of the phone or mobile number are also verified. If the last four digits of three consecutive records are the same and the differences in the first few digits are only within two characters, they are considered the same contact information and merged into one entity. After classification and duplicate removal, all duplicate or nearly duplicate information is centrally marked, and then each marked entity is assigned to the corresponding category, such as applicant information category, address category, and contact information category. To track each entity in subsequent steps, the unique identifier numbers and field sources of these entities are uniformly recorded. When all duplicate removal operations are completed, an application sharing entity set is finally generated.

[0091] According to the set of shared entities in the applications, count the number of occurrences of each entity in different government service applications item by item and mark the category to which the entity belongs. During the execution process, first count the occurrence frequency of each entity category. For example, in the personal information category, count the number of occurrences of the name field and compare it with the same ID number. When the name field and the ID number are the same, it is recorded as the same person, and the total number of occurrences in all government service application forms is recorded. If the ID numbers match but there are three or more characters inconsistent in the name field, this record is excluded and marked in the suspected error list. During the counting process, special attention needs to be paid to whether the ID number appears repeatedly in multiple business types. At the same time, the statistical information of the address entity is distinguished at the prefecture-level city and district / county levels to correspond to a more accurate regional index. If it is found that the same street name appears repeatedly in more than ten application data, this street name is classified into the high-frequency address list. The upper limit of ten is obtained by counting the distribution quantity of government service applications in the past year. After confirming these high-frequency addresses, a secondary comparison will be made in combination with the contact information field. When the same phone number appears in three or more completely duplicate records, this number is included in the high-frequency contact information list. After all entries are counted, the set of shared entity distribution structures is finally obtained.

[0092] Formula: , The benefit of the formula is that it comprehensively processes the cross-frequency, source quantity, matching failure situation, co-occurrence of field positions, and field content changes of entities among multiple government service applications through parameter measurement methods in different dimensions, so as to describe the tightness of the cross-application association relationship of entities in a unified index K, and show the concentration and activity of entities in the cross-application scenario.

[0093] The steps to obtain the parameters are as follows: count the total number of entity categories in the set of shared entity distribution structures, and accumulate all the entity types that have been confirmed to be classified to obtain the final value. This entity type can include personal identity type, address type, contact information type, etc. The determination of each type depends on the previous identification of field names, formats, and possible identifiers. The following gives an acquisition example: After identifying the three major categories of personal information, address information, and contact information in 400 application records and splitting them into 13 sub-categories, these sub-categories are summarized, and the total number of categories is recorded as ETC = 13.

[0094] The steps to obtain the parameters are as follows: count the For the application frequency of class - shared entities, it is necessary to accumulate the number of occurrences of the shared entity in all standardized application data units. Each occurrence is counted as 1 time. If it appears repeatedly, it is incremented by 1. If it is merged into a unique identifier in two application data records, it is only counted as 1 occurrence. These rules are formulated in advance with reference to the process for detecting duplicate entities. The following gives an example: When the class - 1 shared entity "Zhang XX" is identified as the same person in 8 different government service applications, then 。

[0095] The steps to obtain the parameter are to calculate the application cross - count of class - shared entities. It is necessary to detect whether the entity appears with multiple associations within the same application data unit. For example, when the applicant field and the guarantor field both point to the same person, it is regarded as 1 cross. When associated with other fields, each is incremented by 1. Each time a new cross is encountered, it is accumulated in the cross - count of the corresponding entity. The following gives an example: When the class - 2 shared entity "Wang XX's address" is listed as both the enterprise registration address and the contact address in the same application unit, the cross - count is incremented by 2. After synthesizing multiple application units, the AXC value is obtained.

[0096] The steps to obtain the parameter are to record the distribution count of class - shared entities in the sources of government service applications. It is necessary to view the application sources of different business types or different regions, distinguish the occurrence of the entity in each source, and accumulate the source count. This source count can be statistically analyzed separately by application type code and geographical code. The following gives an example: The class - 3 shared entity appears in the business application forms of 5 different provinces and cities, and each province and city belongs to an independent business source, then 。

[0097] The steps to obtain the parameter are to count the number of matching failures of class - shared entities in different standardized application data units. It is necessary to view the occasions where the entity is judged inconsistent during deduplication and classification. If there is a mismatch in field format or partial content conflict, it is recorded as one matching failure. This judgment is based on the similarity threshold when comparing fields such as name, address, or contact information. The following gives an example: When the class - 4 shared entity has 3 cases where it cannot be merged into the same object due to the keyword difference in the name being greater than 0.3, then 。

[0098] The steps to obtain the parameter are to calculate the The number of field positions where class - shared entities co - occur in all application forms. It is necessary to segment and count whether the entity appears in multiple field positions within the same application unit. For example, if it appears in both the applicant and contact information fields in the same form, it is recorded as 2 co - occurrences, and then the results in different units are accumulated. The following gives an acquisition example: When the 5th class - shared entity "Li XX" appears in both the guarantor field and the emergency contact field in a certain application form, the co - occurrence number is incremented by 2. If such a situation occurs in 5 application forms, these times need to be accumulated to obtain 。

[0099] The steps to obtain the parameter are as follows: Record the number of times the field content of the class - shared entity changes. By comparing whether there are differences in the text and numerical information presented by the entity in multiple application units, if the change value includes a slight addition to the address, an update of the last few digits of the contact information, etc., it is regarded as one content change. Different types of changes are accumulated into the same statistical value. The following gives an acquisition example: When the 6th class - shared entity "Zhang XX" has a difference of 2 digits in the last four digits of the phone number in two consecutive applications, it is recorded as 1 change, and when there is an additional street in the address in other applications, it is recorded as 1 change again. The total number of changes is recorded as 2. Therefore 。

[0100] Calculation process:

[0101] In an example, assume the total number of entity classes , for the 1st class of entities, 、 、 、 、 、 , for the 2nd class of entities, 、 、 、 、 、 ;

[0102] First, calculate the part inside the parentheses for the 1st class of entities:

[0103] ;

[0104] , so the above formula becomes:

[0105] ;

[0106] Then, calculate the part inside the parentheses for the 2nd class of entities:

[0107] ;

[0108] , so the above formula becomes:

[0109] ;

[0110] Then add these two parts and multiply by :

[0111] ;

[0112] The result shows that in this example, the cross - application association tightness index is approximately 3.891, indicating that there is a certain degree of association between two types of shared entities in terms of multi - dimensional statistics such as their occurrence frequencies, application cross - over degrees, and field changes. If the result value is relatively large, it means that the cross - application entities show a high degree of crossover and co - occurrence. The smaller the value, the less closely related the different applications are.

[0113] The steps to obtain the multi - application association graph are as follows: Based on the cross - application association tightness index, extract the applicant field, associated address field, and contact information field from multiple standardized application data units, perform entity normalization identification coding and entity type annotation on the content of each field to generate an application entity node set;

[0114] According to the application entity node set, calculate the strength value of the sharing relationship between application entities. The calculation formula is:

[0115] ;

[0116] Among them, is the strength value of the sharing relationship between the th application entity and the th application entity, and are respectively the application frequencies of the th and the th application entities in all standardized application data units, and are the position numbers of the entities in the field structure, and are the lengths of the entity field values, is the number of co - occurrences of the th application entity and the th application entity;

[0117] Based on the strength value of the sharing relationship, construct a graph structure with application entities as nodes and the strength value of the sharing relationship as edge weights, and establish a mapping table of entity edge connection relationships between nodes to generate a multi - application association graph.

[0118] Specifically, based on the cross-application association tightness index, first, the applicant field, associated address field, and contact information field in multiple standardized application data units are sorted out. The cross-application association tightness index obtained from the previous stage is associated with the unique number of each data unit. It is determined that the field types to be processed include personal identity information, specific address descriptions, and contact phone numbers or mobile phone numbers. Then, each of these fields is checked item by item for inconsistent formats or missing values. If it is found that the format does not conform to the standard definition, for example, the number of digits of the ID card number is less than 18, or the district or county name is missing in the address field, it is revised or marked according to the previously established field comparison table. After confirming that all fields can be used for subsequent normalization operations, the field content needs to be segmented item by item. For example, in the address information, the province, city, district or county, and street are divided into multiple segments. For the contact information, only the pure numbers are retained and it is checked whether the number of digits is between 8 and 13. The text information such as the name or legal person name is split into the surname and the rest, and it is checked whether there are spaces or special characters. This can ensure that there is no confusion when assigning unique identification codes to these field contents in the future. When coding, it can be selected to start with the city code or the phone number prefix and add an incrementing serial number at the end to form a unique identifier. For example, for a certain address field, if it matches a certain city, its prefix can be set as the initials abbreviation of that city and the street information can be recorded in the subsequent incrementing numbers. At the same time, the entity is marked as the "address" type. For the applicant field, the last four digits of the ID card number can be combined as a special annotation to distinguish the scenario of the same name but different people. For example, the ID card tail number 1234 of "Wang XX" is encoded as "WANG1234" and marked as "individual entity" in the type annotation. A similar method is used to code the contact information fields. Further, after the coding is completed, repeated verification is carried out. When it is found that the same number or the same name may be assigned multiple codes, the original records pointed to by these codes are compared. If the similarity is above 0.8 (this 0.8 is obtained by analyzing the repetition rules of the first 7 digits and the last 4 digits of mobile phone numbers in more than two thousand contact information data last year), they are merged into the same entity code and marked as "high-frequency common entity" for subsequent tracking. After the above processing steps, all field contents are assigned normalized identification codes and clear entity type annotations, and finally an application entity node set is generated.

[0119] Formula: , The benefit of the formula is that through the multiplication and division operations of multiple factors such as the frequency of each entity's appearance, the field position number, the number of co-occurrences, and the difference in field value lengths, the shared connection strength between different entities is comprehensively represented. Functions such as logarithms and square roots are used to weaken the influence of overly extreme values and more closely associate entities with similar or identical content numerically.

[0120] The acquisition steps are as follows: In all standardized application data units, count the number of occurrences of the th applicant entity. Each time the entity is confirmed to appear in any data unit, it is counted as 1 time. If it is merged due to similar fields in a unit, it is still regarded as one occurrence. The entire statistical process requires matching the entity ID and performing duplicate detection. For example, in 200 government applications, records with the last four digits of the ID card of "Zhang Moumou" or "Zhang XX" being the same are unified and merged. If this person appears 37 times in all 200 records, then it can be recorded as .

[0121] The acquisition steps are as follows: In the same set of standardized application data units, count the number of occurrences of the th applicant entity. The method is the same as . Compare the entity's independent ID with information such as name and address, and accumulate all occurrences. If there are slight differences in the name writing or the mobile phone number has a separator, the previously set similarity determination and duplicate removal still need to be followed. If a certain entity appears 50 times in 300 records, then record it as .

[0122] The acquisition steps are as follows: Identify the position number of the th applicant entity in the field structure. The position number needs to be arranged according to the inherent order of the fields in the application form. For example, the applicant field is numbered 1, the guarantor field is numbered 2, the contact information field is numbered 3, the address field is numbered 4, and so on. If it is found that the th applicant entity is the guarantor of a certain application unit, then .

[0123] The acquisition steps are as follows: Following the same idea as , number and count the positions where the th applicant entity appears in all application forms. When counting, if the entity mainly appears in the address field, the number may be 4. If it appears in the applicant field in more scenarios, the number is 1. Finally, confirm the value according to the appearance frequency. If the information of a certain enterprise or institution is located in the "associated address" field six times in multiple application forms, then .

[0124] The acquisition steps are as follows: Count the For the field value length of an applicant entity, first confirm which category the entity belongs to. If it is a name category, count the number of characters. If it is an ID card number category, it is 18. If it is an enterprise name, look at the total number of characters in its text in the application form. If it is an address, count according to the completeness of the specific address. Select the text length of the most complete occurrence of the same entity among multiple occurrences as , for example, the address field length in a certain city is counted as 18 characters.

[0125] The acquisition steps for are as follows: In the th applicant entity, also calculate the number of characters according to its most complete occurrence form. If the entity appears as "XX Street, a certain place" and "XX District, XX Street, a certain place" in several applications, select the form with the larger number of characters for counting. Once it is determined to be 21 characters, then .

[0126] The acquisition steps for are as follows: Count the number of times the th applicant entity and the th applicant entity appear together. It is necessary to scan all standardized application data units. When both entities appear in the same application, count 1 time. If they are mentioned in multiple locations in the same application form, it is also only counted as 1 time. If it is found after scanning that these two entities are recognized together in 10 applications, then .

[0127] Calculation process:

[0128] In an actual example, for example, the parameter values of the th entity and the th entity are as follows:

[0129] , , , , , ,

[0130] First, calculate the numerator part :

[0131] ;

[0132] ;

[0133] ;

[0134] ;

[0135] Then, calculate the denominator part :

[0136] ;

[0137] ;

[0138] ;

[0139] ;

[0140] Therefore, the entire formula gives:

[0141] ;

[0142] This result indicates that when the occurrence times of both entities are relatively high, the field position numbers are far apart, and the number of times they co-occur in the application is not too low, the strength value of their sharing relationship will show a relatively high value. When it is higher than 25, it can be regarded as being closely related in government application data. If it is less than 5, it means that the association between the two in the data is not obvious.

[0143] Based on the sharing relationship strength value, first traverse all the strength calculation results between entities in the previous step and record them item by item. For the same entity node and other entity nodes, if there is a sharing relationship reflected in multiple application forms, the calculated strength values are retrieved in order according to the recording sequence for comparison, and the values higher than 30 are selected as tight connections. Here, 30 is obtained by adding three times the standard deviation to the statistical mean of the past ten thousand entity pair relationship data. This can ensure that it is only judged as a tight connection under truly high-strength conditions. At the same time, the strength values lower than 10 are marked as loose connections in the graph structure. This 10 is set by the lower quartile found after summarizing the same number of entity pair relationships. Whenever a strength result is inserted, an edge is established between the two according to the entity node numbers, and the weight of the edge is set to the corresponding strength value. If it is found that the strength value of the entity pair already exists in other records, an overwrite is performed to retain the latest calculation result. After all the sharing relationship strength values between entity nodes are inserted or updated, the system will generate a graph structure with entity IDs as nodes and strength values as edge weights. In this structure, querying any entity node can quickly obtain the connection weights and connection relationships between it and other nodes. Further, a mapping table between nodes can be exported, including information such as the starting entity identifier, ending entity identifier, edge weight, and connection direction. After uniformly organizing the records of these nodes and edges, a multi-application association graph is finally generated.

[0144] The steps for obtaining the comprehensive review risk prompt are as follows: Based on the multi-application association graph, extract the bidirectional paths of all application entity nodes from the graph structure, and retrieve the single-application compliance status determination results, path lengths, out-degrees of the path end nodes, number of fields of the path start nodes, and the number of intermediate nodes passed by each path to generate a multi-application path attribute set;

[0145] According to the multi-application path attribute set, calculate the risk score of the corresponding entity node. The calculation formula is:

[0146] ;

[0147] where, is the risk score of the th application entity node, is the single-application compliance status determination result of the end point of the th path. Compliance is 1 and non-compliance is 0, is the out-degree of the path end node, is the path length, is the number of fields of the path start node, is the number of intermediate nodes of the path, is the total number of entity fields of the path end node in the multi-application association graph, is the th number of paths involved by the application entity node;

[0148] Based on the risk score, divide all application entity nodes into risk level intervals according to the risk score, and mark the risk level color code and risk warning node number in each application path to generate a comprehensive review risk prompt.

[0149] Specifically, based on the multi-application association graph, first read the bidirectional paths of the corresponding entity nodes from the graph structure one by one. Each path can be traced back to the edge weights between the nodes collected previously, and compared with the compliance or non-compliance results recorded in the compliance status determination link of a single application. To ensure the accuracy of path extraction, it is necessary to scan the out-degree and in-degree of each entity node, and select the node combination that can form a complete bidirectional link. When analyzing the path, it can be divided into short paths and long paths according to the length of the path. The path length is defined as the number of intermediate nodes plus the number of steps taken by the two endpoints. If the length of a path exceeds 5, this 5 is the threshold confirmed after statistical analysis of hundreds of graph data in the previous stage. It is included in the long path group that needs to be recorded. When retrieving each path, the out-degree of the node at the end of the path will be extracted and attached to the intermediate result. The out-degree value is mainly obtained by the node connection when aggregating the graph data. If It is found that the out-degree value is above 8. The 8 here is determined as the dividing line exceeding the average plus twice the variance by analyzing the out-degree distribution of three thousand entity nodes. It is temporarily determined to be a high out-degree node. Next, it is necessary to retrieve the number of fields of the starting node of the path, in which it will be checked whether the total number of field information allocated to the entity node in the previous standardization stage exceeds 10. This 10 is a constant statistical value based on the items contained in common personal or corporate information. If it exceeds 10, it is classified as a node with more information. In addition, the number of intermediate nodes in the path must be read, and the additional nodes passed by each link must be counted. These nodes are marked in a list and combined with the compliance status of the path endpoint. If there is a non-compliant judgment result at the end node of the path, the "CR" item will be marked as 0 in the attribute set. If it complies with the rules, it will be marked as 1. When all entries are processed, the path length, the end point out-degree, the number of fields of the starting node, and the number of intermediate nodes in the path can be merged into a multi-application path attribute set.

[0150] formula: The benefit of the formula is that it combines the information such as the compliance status of the path endpoint, the node out-degree, the path length, and the difference in the number of entity fields through the combination of different indicator items, providing a relatively comprehensive risk quantification result. Combined with the individual situation of each path, the overall risk score for the application entity node is obtained, which is convenient for subsequent hierarchical management and control.

[0151] The steps to obtain are: in the The end node of a path queries the compliance status determination result of a single application. When the compliance determination value associated with this node is compliant, it is recorded as 1; when the determination value is non-compliant, it is recorded as 0. These determination results come from the previous single-application verification process, which can be determined after completing operations such as business type matching, numerical range comparison, and geographical mapping. If in two different paths, for the same node, one compliance result is 1 and the other is 0, they should be recorded separately in the corresponding path information. The following gives an example of obtaining: If it is found through inspection that all field values of the corresponding single application of the end node of a certain path are in line with the clause requirements, then 。

[0152] The obtaining steps are as follows: Count the out-degree of the end node of the path. The out-degree represents the number of outgoing edges connected by this node in the multi-application association graph. Check each connection line between the node and other nodes one by one. If a total of 6 edges are connected, then 。

[0153] The obtaining steps are as follows: Record the length of the th path. If the path contains several intermediate nodes from the start point to the end point, the total number of nodes - 1 can be recorded as the path length, or directly count the number of edges contained in the path. For example, when starting from point A and passing through 3 intermediate nodes to reach point Z, the path length can be recorded as 4. The meaning of this 4 is the total number of edges along the way. The following gives an example of obtaining: In a specific link, there are 2 intermediate nodes C and D between the start point B and the end point E. Therefore, there are 3 edges in the path. If the project stipulates that the path length = the number of edges, then 。

[0154] The obtaining steps are as follows: Read the number of fields of the start node of the path. In the previous standardization stage, each entity node has saved the number of field description entries it owns. If a certain enterprise node contains a total of 12 fields such as name, business license number, legal person information, and business scope, then 。

[0155] The obtaining steps are as follows: Count the number of intermediate nodes passed through in the th path, excluding the start node and the end node. For example, in the same path from A to F passing through two nodes B and C, the number of intermediate nodes is 2. If a certain path does not contain intermediate nodes, then 。

[0156] The obtaining steps are as follows: Obtain the total number of entity fields of the end node of the path in the multi-application association graph. This value is the same as Similarly, except that the query is for the end node. In the standardization stage, each node is identified with the field items it owns. If the end node contains a total of 8 fields such as ID card number, mobile phone number, landline phone, and home address, then .

[0157] The acquisition steps of are: count the number of paths involved in the th application entity node, that is, check the number of times this node appears as a starting point or an end point in the graph structure and forms a complete link. If a node can be both a starting point and an end point, all the paths it involves need to be counted. When scanning item by item in the graph data, as long as the starting point = this node or the end point = this node is shown in the path record, a count is made, and the sum is obtained until the total value. The following is an acquisition example: When a certain node appears 5 paths with it as the starting point and 3 paths with it as the end point in the graph, then .

[0158] Calculation process:

[0159] Give a specific calculation example. Let an entity node be involved in a total of 3 paths. Therefore , number these 3 paths as k = 1, k = 2, k = 3 in sequence, and given the corresponding parameters as follows:

[0160] Path 1: , , , , , ;

[0161] Path 2: , , , , , ;

[0162] Path 3: , , , , , ;

[0163] First calculate the items of Path 1:

[0164] ;

[0165] ;

[0166] ;

[0167] In summary, the total of Path 1 items:

[0168] ;

[0169] Calculate Path 2 items:

[0170] ;

[0171] ;

[0172] ;

[0173] Total of Path 2 items:

[0174] ;

[0175] Calculate Path 3 items:

[0176] ;

[0177] ;

[0178] ;

[0179] Total of Path 3 items:

[0180] ;

[0181] Add the result values of the three paths:

[0182] ;

[0183] Then perform square root extraction and divide by :

[0184] ;

[0185] This result indicates that in this example, the entity node appears at the end of 1 compliant node and 2 non-compliant nodes in the 3 involved paths. At the same time, factors such as out-degree, path length, number of fields, and number of intermediate nodes jointly affect the final risk score. The result value of 0.90 indicates relatively moderate. If 1.5 is set as the higher risk threshold during grading, it does not reach the high-risk range. If it is lower than 0.5, it can be regarded as having a lower risk.

[0186] Based on the risk score, the score of each entity node is compared one by one with the previously set upper and lower limits of the interval. If the score of a certain node exceeds 1.5, where 1.5 is obtained by calculating the average value of 1.0 for all previous node score samples and adding a safety factor of 0.5, it is marked as a high-risk. If the score is between 0.5 and 1.5, it is classified as a medium-risk. If the score is below 0.5, it is regarded as a low-risk node. During path retrieval, high-risk nodes are marked in red, medium-risk nodes are marked in yellow, and low-risk nodes are marked in green, and the corresponding identifiers are filled in the attribute table. If it is found that there are multiple high-risk nodes in a path, the path will be bolded in the graph display interface, and at the same time, all node numbers of this path will be recorded in the warning list. When all nodes are traversed and marked, a comprehensive review risk prompt is finally generated.

Claims

1. An AI automatic review system for government affairs services, characterized in that, The system includes: An application information structuring module, which, based on government service application materials, identifies and extracts elements such as applicant type, business type, region, amount, contact information, address, and guarantor in forms and text fields, maps the elements to preset data fields, establishes internal association identifiers, and establishes a standardized application data unit; A rule compliance verification module, which, based on the standardized application data unit, retrieves the rule condition clauses in the database that match the application business type and qualifications, calculates the rule application complexity factor according to the number of matching clauses and the nesting level, and based on the rule application complexity factor, compares the application data with the clause constraint values, time limits, and regional scopes item by item to determine the compliance status of the attributes and generate a single application compliance status determination; An associated network construction module, which, based on multiple standardized application data units, searches for applicant, associated address, and contact information entities shared among different government service applications, calculates the cross-application association tightness index according to the number and categories of shared entities, and based on the cross-application association tightness index, constructs a graph structure with application entities as nodes and sharing relationships as edges to establish a multi-application association graph; An abnormal pattern recognition and audit recommendation module, which, based on the multi-application association graph, calculates a risk score in combination with the single application compliance status determination of the associated applications and generates a comprehensive audit risk prompt; The steps for obtaining the rule application complexity factor are as follows: Based on the application business type field and application qualification field in the standardized application data unit, extract the original field values and match them with the field definition library for field content normalization, complete the missing fields, and unify the field formats and semantic expression methods to obtain a combined set of normalized application business type fields and application qualification fields; According to the combined set of normalized application business type fields and application qualification fields, sequentially retrieve the rule condition clauses in the database that match the business type, filter the clauses that match the qualification fields and the combined set, and extract the nesting level, number of logical judgments, clause reference times, clause activation period coverage, clause applicable region overlap degree, and call frequency of each clause to generate a structured information set of rule condition clause matches; Based on the structured information set of rule condition clause matches, calculate the rule application complexity factor. The calculation formula is: ; Among them, is the complexity factor of rule application, is the number of logical judgments of the th clause, is the nesting level of the th clause, is the regional overlap degree of the th clause, is the clause reference times of the th clause, is the historical call frequency of the th clause, is the number of clauses in the structured information set that match the rule condition clauses.

2. The government affairs business AI automatic review system according to claim 1, characterized in that, The steps for obtaining the standardized application data unit are as follows: Based on government service application materials, scan and identify the form field content in the materials item by item, perform field category identification and matching, and perform text field content splitting, and extract elements such as applicant type, business type, region, amount, contact information, address, and guarantor one by one to generate a set of original elements of government service application materials; Based on the set of original elements of government service application materials, perform field mapping matching verification one by one, and perform structure format conversion and field reconstruction on the successfully matched field content to generate a set of standardized mapping fields of government service application materials; Based on the standardized mapping field set of the government service application materials, internal association matching is performed among the standardized mapping fields, and internal association relationship identifiers are established according to the matching results. Index construction and relationship binding of the standardized mapping fields are carried out to generate standardized application data units.

3. The government affairs business AI automatic review system according to claim 1, characterized in that, The obtaining steps of the single-application compliance status determination are as follows: Based on the rule application complexity factor, the application business type field values and clause constraint values in the standardized application data units are called item by item to perform conditional threshold comparison, field numerical range inspection, and timeliness matching judgment between the field values and the constraint values, forming a set of intermediate results of the comparison between the application data and the clause constraints. According to the set of intermediate results of the comparison between the application data and the clause constraints, the geographical area field values in the standardized application data units are obtained, and item-by-item mapping of the application geographical area field and the applicable geographical area range of the rule clauses is performed to screen the clauses with successful geographical area mapping matches, generating a set of matching results between the application data and the geographical area range. Based on the set of matching results between the application data and the geographical area range, the compliance situation of the attributes of each application data is determined, and the single-application compliance status determination is generated.

4. The government affairs business AI automatic review system according to claim 1, characterized in that, The obtaining steps of the cross-application association tightness index are as follows: Based on multiple standardized application data units, the applicant fields, associated address fields, and contact information fields in each standardized application data unit are extracted, and entity classification matching and deduplication processing are performed on all application data according to the field types, generating a set of application shared entities. According to the set of application shared entities, the occurrence times of the entities in different government service applications are counted item by item and the entity category to which they belong is marked, and the repeated distribution information of each type of shared entity in each standardized application data unit is recorded, obtaining a set of shared entity distribution structures. Based on the set of shared entity distribution structures, the cross-application association tightness index is calculated.

5. The government affairs business AI automatic review system according to claim 1, characterized in that, The obtaining steps of the multi-application association graph are as follows: Based on the cross-application association tightness index, the applicant fields, associated address fields, and contact information fields in multiple standardized application data units are extracted, and entity normalization identification coding and entity type annotation are performed on the content of each field, generating a set of application entity nodes. According to the set of application entity nodes, the shared relationship strength values between the application entities are calculated. Based on the shared relationship strength values, a graph structure with application entities as nodes and shared relationship strength values as edge weights is constructed, and a mapping table of entity edge connection relationships between the nodes is established, generating a multi-application association graph.

6. The government affairs service AI automatic review system according to claim 1, characterized in that The obtaining steps of the comprehensive review risk prompt are as follows: Based on the multi-application association graph, all bidirectional paths of the application entity nodes are extracted from the graph structure, and the single-application compliance status determination results, path lengths, out-degrees of the path end nodes, number of fields of the path start nodes, and number of intermediate nodes passed by the paths corresponding to each path are retrieved, generating a set of multi-application path attributes. According to the set of multi-application path attributes, the risk scores of the corresponding entity nodes are calculated. Based on the risk scores, all applicant entity nodes are divided into risk level intervals according to the risk scores, and the risk level color coding and risk warning node numbers are marked in each application path to generate a comprehensive review risk prompt.

7. The government affairs service AI automatic review method of the government affairs service AI automatic review system according to any one of claims 1-6, characterized in that, It includes the following steps: Based on the government service application materials, identify and extract the applicant type, business type, region, amount, contact information, address, and guarantor elements in the form and text fields, attribute the elements to the preset data fields through field mapping, and establish the association identifiers between the elements to generate a standardized application data unit; Based on the standardized application data unit, retrieve the database, match the rule condition clauses of the application business type and qualifications, calculate according to the number of clause matches and nested levels to obtain the rule application complexity factor, and based on the rule application complexity factor, compare each application data with the clause constraint values, time limit, and regional scope to determine whether each attribute meets the requirements and generate the compliance status determination of a single application; Based on multiple standardized application data units, find the applicants, addresses, and contact entity shared across applications, calculate the relationship strength between the number of shared entities and categories to obtain the cross-application association tightness index, and based on the cross-application association tightness index, construct a graph structure with application entities as nodes and shared relationships as edges to generate a multi-application association map; Based on the multi-application association map, combined with the compliance status of a single application for each application, calculate the risk score, evaluate the risk across applications, and generate a review risk prompt; Analyze the comprehensive review risk prompt, refine the review opinions, and provide guidance for manual review and automatic decision-making to obtain review suggestions.

Citation Information

Patent Citations

  • Contract intelligent auditing method and device, computer device and storage medium

    CN109886845A

  • Audit rule processing method and device based on machine learning and computer equipment

    CN111813399A