A large-scale master data automatic approval error correction system based on an agent
By using an agent-based large-scale master data automatic approval and error correction system, dynamic clustering and feature association analysis are employed to solve the problem of insufficient approval accuracy in existing automated approval systems, thereby achieving automation and improved accuracy in intelligent decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-10
AI Technical Summary
Existing automated approval systems rely on manual operation, lack context awareness in approval conclusions, fail to detect group anomalies, and have insufficient approval accuracy.
A large-scale master data automatic approval and error correction system based on intelligent agents is adopted, including an approval data access verification module, a multi-category data dynamic clustering module, a multi-data feature correlation deviation analysis module, a multi-type approval rule matching module, and a multi-path approval error correction derivation and verification module. Through dynamic clustering and feature correlation analysis, different levels of approval processes are automatically triggered, and the most suitable approval rules and error correction strategies are applied.
It significantly improves the accuracy of approval decisions, automates intelligent decision-making, and enhances the efficiency and accuracy of data approval.
Smart Images

Figure CN121388529B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent data, specifically to a large-scale master data automatic approval and error correction system based on intelligent agents. Background Technology
[0002] In many fields such as finance, insurance, supply chain management, and personnel reimbursement, there is a large amount of structured or semi-structured data that requires approval. Data approval is a key step in data preprocessing in statistics. It refers to the review and verification of the completeness and accuracy of raw data, aiming to ensure data quality and improve analytical efficiency. After collecting statistical data through various channels, the first step is to process and organize this data to make it systematic and organized to meet the needs of analysis. Data organization typically includes several aspects such as data preprocessing, classification or grouping, and summarization. It is a necessary step before statistical analysis. Data preprocessing is a preliminary step before data grouping and organization, and includes data review and screening, data sorting, etc.
[0003] Traditional data approval processes rely heavily on manual operation. Approvers need to check the data one by one based on written rules and regulations and experience. Existing automated approval systems, based on rules, mostly focus on making "absolute" judgments on individual data instances. However, without comparison with similar categories, the approval conclusions may be inaccurate. The approved data lacks context awareness and cannot detect group anomalies.
[0004] This application aims to preprocess and dynamically cluster approval data input from the approval data source to construct data-driven dynamic group profiles of different categories. Individual data segments are placed in the distribution of their respective group categories for correlation analysis. Based on the degree of deviation of individual data segments from the group benchmark, different levels of approval processes are automatically triggered. At the same time, by classifying data categories, the most suitable approval rules and error correction strategies are applied to different categories of data, prioritizing the error correction scheme with the highest confidence, which significantly improves the accuracy of approval decisions and realizes intelligent decision automation. Summary of the Invention
[0005] The purpose of this invention is to provide a large-scale master data automatic approval and error correction system based on intelligent agents to solve the problems in the prior art.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A large-scale master data automatic approval and error correction system based on intelligent agents is disclosed. The system includes an approval data access and verification module, a multi-category data dynamic clustering module, a multi-data feature correlation deviation analysis module, a multi-type approval rule matching and correspondence module, a multi-path approval error correction derivation and verification module, and a manual approval optimization platform. The approval data access and verification module, the multi-category data dynamic clustering module, the multi-data feature correlation deviation analysis module, the multi-type approval rule matching and correspondence module, and the multi-path approval error correction derivation and verification module are sequentially connected to the manual approval optimization platform.
[0008] The approval data access verification module is used to preprocess the approval data input from the approval data source, perform preliminary compliance verification of the data source by field, and perform standardized verification of the multi-field format of the input data.
[0009] The multi-category data dynamic clustering module is used to perform multi-dimensional comparison of key fields in the text of different pending approval data, build a dynamic clustering model, dynamically cluster different categories of data, automatically classify different input data into specific data categories, build a pre-similar dataset, perform similarity analysis based on the input data within different pre-similar datasets, perform dual matching of high-frequency field overlap and intent relevance on the input data within the pre-similar datasets, analyze and build similar datasets, and build a group data feature profile of similar datasets based on different data categories;
[0010] The multi-data feature association deviation analysis module is used to associate features of image data and special data of each input data in the same category dataset. It analyzes and filters out the features of image data and special data of different input data and measures the degree to which the input data deviates from the group mean in terms of semantic information features.
[0011] The multi-type approval rule matching module matches the corresponding approval rules based on the data of the same category, performs intelligent verification on the data of the same category in sequence according to the matched approval rules, performs random secondary approval feasibility verification based on the data approval results, and provides feedback on abnormal approval data.
[0012] The multi-path approval error correction derivation and verification module locates erroneous data based on different data approval correction opinions, provides error correction prompts based on the data error type, performs confidence analysis on each of the multiple error correction prompts, and uploads a backup after user feedback confirmation on the error correction prompts after confidence verification.
[0013] Further configuration: The approval data access verification module includes a multi-source data preprocessing submodule and a data format standardization verification submodule. The multi-source data preprocessing submodule acquires raw data from different data sources, cleans and preprocesses the raw data, and verifies the input data format, data sending time, and data sending user. It removes data with obvious anomalies such as missing data, format errors, abnormal sending time, and unclear sending user. The data format standardization verification submodule acquires the preprocessed data and divides different data according to different format fields. Format fields include text data, image data, and special data. Special data includes numeric data, alphabetic data, and symbolic data. It screens different fields according to different standardization format rules. The standardization format rules are uploaded manually. It screens the non-standardization rate of different fields in each input data. When the non-standardization rate of a certain field exceeds a set threshold, it obtains the average proportion of the total non-standardization rate of the input data. When the average proportion of the non-standardization rate of the input data exceeds the set threshold, the input data is sent back to the sending user for correction.
[0014] Further configuration: The multi-category data dynamic clustering module includes a multi-dimensional comparison and discrimination module for similar datasets and a pre-similar data feature similarity analysis clustering sub-module. The multi-dimensional comparison and discrimination module for similar datasets acquires pre-processed data to be approved, extracts arbitrary keywords from the text data within the input data to be approved, marks key data features, performs pre-comparison based on the marked key data features of different input data, analyzes the overlap of key data features based on different input data, classifies input data with key data feature overlap exceeding a set threshold, and marks them as pre-similar input datasets, screens high-frequency words with pre-similar input datasets, performs intent relevance judgment based on the screened high-frequency words, and screens the intent relevance between each unclassified input data and each pre-similar input dataset. When the intent relevance between the unclassified input data and each pre-similar input dataset is greater than a set threshold, the input data is classified into the pre-similar input dataset.
[0015] Further settings: The pre-similar data feature similarity analysis clustering submodule obtains each pre-similar dataset, analyzes the similarity of the input data within each pre-similar dataset, and sets the similarity score of the input data within each pre-similar dataset as follows. The input data within each pre-class dataset is matched against the high-frequency words in the pre-class dataset using field overlap matching. The overlap rate between the input data within each pre-class dataset and the high-frequency words is set to a certain percentage. The process involves filtering out data fields within each pre-class dataset that overlap with high-frequency word fields. Then, it performs intent relevance matching between the remaining fields of the input data within each pre-class dataset and the intents related to each high-frequency word in the pre-class dataset. The intent relevance between the remaining fields of the input data within each pre-class dataset and the intents related to each high-frequency word in the pre-class dataset is set to [value missing]. A multi-factor weighted model was used to perform similarity analysis between the input data within different pre-class datasets and the dataset of the same category. The weighting of the high-frequency words in the input data and the pre-class input datasets was set to account for overlap in the matching of fields. The input data is matched with high-frequency words in a pre-classified input dataset based on intent association, with the weighting ratio being... Set the similarity score threshold for input data within each pre-class dataset to be [value]. According to the formula:
[0016]
[0017] The similarity score of each input data set within a pre-class dataset is calculated to determine the similarity between the input data and the pre-class input dataset. If the above formula is satisfied, it means that the input data belongs to the same category within the pre-class input dataset. When the similarity score of each input data belongs to the same category... If the above formula is not met, the input data is removed and reclassified until the similarity score between the input data and the corresponding pre-classified dataset meets the above threshold. Each input data within each pre-classified dataset is screened, and if the similarity score between the input data and the corresponding pre-classified dataset meets the above threshold, it is re-labeled. The pre-classified dataset corresponding to this type of input data is labeled as a class dataset. Based on the category of each class dataset, a group text data feature profile of the class dataset is constructed.
[0018] Further configuration: The multi-data feature association deviation analysis module includes a data feature dynamic association identification screening module and an association data deviation value quantification analysis submodule. The data feature dynamic association identification screening module is used to acquire image data and special data of each input data in the same dataset, pre-identify text regions and image regions in the image data, extract text information within the text regions, convert it into text semantic vectors through a natural language processing model, extract the image part of the image region for content recognition, and determine whether the recognized content of the image region matches the text information of the text region. If they match, extract the group text data feature profile within the same dataset, compare the text semantic information of the text region with the group text data feature profile within the same dataset, analyze the correlation between the image data of each input data and the group text data feature profile within the same dataset, and filter out data in the same dataset where the correlation between the image data of each input data and the group text data feature profile within the same dataset is less than a set threshold.
[0019] Obtain the special data of each input data in the same dataset, extract the statistical features, numerical distribution features and syntactic structure features of the special data respectively, compare the features of the special data of the input data in each dataset, and filter out the input data corresponding to the obvious abnormal special data features.
[0020] The input data corresponding to the screening of image data and the feature profiles of the group text data within the same dataset that have a correlation degree of less than a set threshold and the presence of obvious abnormal special data features are statistically analyzed. The input data is sent to the manual approval and optimization platform for review. The same dataset after the screening operation is obtained, and the same dataset is marked three times to become the verification dataset.
[0021] Further settings: The correlation data deviation quantification analysis submodule acquires each input data within the same verification dataset, obtains the corresponding group text data feature profile of the same verification dataset, extracts each core feature within the corresponding group text data feature profile, extracts the semantic information of each core feature, compares the text data semantic information features of the input data and the text semantic information features of the text region in the image data with any core feature for single feature deviation, the deviation judgment rule is uploaded by the user, analyzes the average deviation of the text data semantic information features of each input data and the text semantic information features of the text region in the image data with the semantic information features of each core feature, combines the average deviation of the text data semantic information features of each input data and the text semantic information features of the text region in the image data with the semantic information features of each core feature, calculates the comprehensive deviation score of the input data, and sets a reasonable range for the comprehensive deviation. Set the overall deviation score of a certain input data as Set the deviation quantization value for different input data. According to the formula:
[0022]
[0023] Obtain the deviation quantization values of different input data within the same dataset for verification, and divide the different input data within the same dataset for verification based on the different deviation quantization values.
[0024] Further configuration: The multi-type approval rule matching module includes a multi-dataset approval rule matching sub-module, an approval strategy feasibility call verification sub-module, and an abnormal approval reporting feedback sub-module. The multi-dataset approval rule matching sub-module is used to obtain the feature profile of the group text data of the same type of dataset for verification, obtain the uploaded approval rule library data, match different approval rules according to the feature profile of the group text data, and approve the input data within each verification dataset according to different approval rules to generate approval results. If the approval result of a certain input data fails, the final approval opinion is generated according to the corresponding approval rule. The approval opinion includes at least one specific error type and misalignment location information, and is sent to the multi-path approval error correction derivation and verification module.
[0025] The feasibility verification submodule for the approval strategy calls upon the user to obtain input data with deviation quantization values of 1 and 0 from the same verification dataset. It then extracts input data with deviation quantization values of 1 and 0 from each verification dataset for secondary approval verification. The amount of input data with deviation quantization value of 1 is set to be... The input data volume is 1, with a deviation quantization value of 1. , The system screens the first and second approval results of input data with deviation values of 1 and 0 in each sampled dataset of the same type of data. It verifies whether the first and second approval results of the same input data are consistent. If the approval results are consistent, it marks that the verification strategy for matching the same type of data is effective. If the approval results are inconsistent, it screens the inconsistent data and marks them as abnormal. The abnormal approval reporting feedback submodule sends the abnormally marked input data to the manual approval optimization platform for manual approval.
[0026] Further setup: The multi-path approval error correction derivation and verification module includes a multi-data approval correction and error correction data marking submodule, a correction value confidence analysis verification submodule, and a correction result manual feedback optimization confirmation submodule. The multi-data approval correction and error correction data marking submodule obtains the final approval opinions based on several input data within each verification similar dataset. Based on the final approval opinions of different input data, including error type and misalignment location information, it determines the error type of the input data. Error types include format errors, logical errors, partial data errors, and partial information missing. For format errors and logical errors, it corrects them according to the correction database uploaded by the manual approval optimization platform. It summarizes the corrected input data and sends the corrected input data to the corresponding approval rules in the approval rule base for secondary approval. When the approval is passed, it replaces the original input data within the verification similar dataset with the corrected input data to form a new verification similar dataset. For input data with partial data errors and partial information missing, it obtains the historical sending data of the user who sent the input data and finds the K historical sending data records that are most similar to the target input data, using the mode of their corresponding fields as candidate missing data.
[0027] The modified value confidence analysis verification submodule obtains candidate missing data from the target input data, retrieves the historical data path of the candidate missing data, analyzes the number of times it appears in the historical data, and sets the number of times it appears in the historical data as [value missing]. The baseline confidence level for its candidate missing data is set to 1. , = The candidate missing data is mapped to the context of the target input data, and the matching degree between the candidate missing data and the target input data context is analyzed. The number of context fields in the target input data that are logically related to the candidate missing data is set to 1. The contextual fit between the candidate missing data and the target input data is set to 1. According to the formula:
[0028]
[0029] In the above formula, For candidate missing data values, After filling in the missing candidate data in the target input data, the new input data is formed with the first... The values of the context fields This is a consistency indicator function used to determine the collected candidate missing data values. With any context field Are they logically similar when collecting candidate missing data values? and If the candidate missing data values are similar and logically consistent, the function output will be 1. and If they are different and logically contradictory, the function output will be 0;
[0030] Analyze the final confidence level of the candidate missing data, and set the final confidence level of the candidate missing data as . According to the formula:
[0031]
[0032] in, and These are the base confidence level of manually set candidate missing data and the configuration weight coefficients for the contextual fit between candidate missing data and target input data, respectively. ,when If the value exceeds a set threshold, the candidate missing data is substituted into the corresponding input data to replace the erroneous fields and fill in the missing information. It is then marked as pre-corrected data. The pre-corrected data is sent to the data sending user for confirmation through the manual feedback optimization confirmation submodule. After confirmation, the input data in the corresponding similar verification dataset is replaced to form a new similar verification dataset. The new similar verification dataset is then sent to the manual approval optimization platform for backup.
[0033] Compared with existing technologies, the beneficial effects of this invention are: it aims to preprocess and dynamically cluster the approval data input from the approval data source to construct data-driven dynamic group profiles of different categories, place individual data segments in the distribution of their respective group categories for correlation analysis, and automatically trigger different levels of approval processes based on the degree of deviation between individual data segments and the group benchmark. At the same time, through data classification, it applies the most suitable approval rules and error correction strategies to different categories of data, prioritizing the error correction scheme with the highest confidence, which significantly improves the accuracy of approval decisions and realizes intelligent decision automation. Attached Figure Description
[0034] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0035] Figure 1 This is a schematic diagram of the specific structure of a large-scale master data automatic approval and error correction system based on intelligent agents according to the present invention;
[0036] Figure 2 This is a schematic diagram of the module structure of a large-scale master data automatic approval and error correction system based on intelligent agents according to the present invention. Figure 1 ;
[0037] Figure 3 This is a schematic diagram of the module structure of a large-scale master data automatic approval and error correction system based on intelligent agents according to the present invention. Figure 2 ;
[0038] Figure 4 This is a schematic diagram of the module structure of a large-scale master data automatic approval and error correction system based on intelligent agents according to the present invention. Figure 3 ;
[0039] Figure 5 This is a schematic diagram of the module structure of a large-scale master data automatic approval and error correction system based on intelligent agents according to the present invention. Figure 4 ;
[0040] Figure 6 This is a schematic diagram of the module structure of a large-scale master data automatic approval and error correction system based on intelligent agents according to the present invention. Figure 5 . Detailed Implementation
[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] Please see Figures 1-6 In this embodiment of the invention, a large-scale master data automatic approval and error correction system based on intelligent agents is provided. The system includes an approval data access verification module, a multi-category data dynamic clustering module, a multi-data feature correlation deviation analysis module, a multi-type approval rule matching module, a multi-path approval error correction derivation and verification module, and a manual approval optimization platform. The approval data access verification module, the multi-category data dynamic clustering module, the multi-data feature correlation deviation analysis module, the multi-type approval rule matching module, and the multi-path approval error correction derivation and verification module are sequentially connected to the manual approval optimization platform.
[0043] The approval data access verification module is used to preprocess the approval data input from the approval data source, perform preliminary compliance verification of the data source by field, and perform standardized verification of the multi-field format of the input data.
[0044] according to Figure 2Furthermore, it needs to be explained in more detail that the approval data access verification module includes a multi-source data preprocessing submodule and a data format standardization verification submodule. The multi-source data preprocessing submodule obtains raw data sent from different data sources, cleans and preprocesses the raw data, and verifies the input data format, data sending time, and data sending user. It removes data with obvious anomalies such as missing data, format errors, abnormal sending time, and unclear sending user. The data format standardization verification submodule obtains the preprocessed data and divides different data according to different format fields. Format fields include text data, image data, and special data. Special data includes numeric data, alphabetic data, and symbolic data. It screens different fields according to different standardization format rules. The standardization format rules are uploaded manually. It screens the non-standardization rate of different fields in each input data. When the non-standardization rate of a certain field exceeds a set threshold, it obtains the average proportion of the total non-standardization rate of the input data. When the average proportion of the non-standardization rate of the input data exceeds the set threshold, the input data is sent back to the sending user for correction.
[0045] The multi-category data dynamic clustering module is used to perform multi-dimensional comparison of key fields in the text of different pending approval data, build a dynamic clustering model, dynamically cluster different categories of data, automatically classify different input data into specific data categories, build a pre-similar dataset, perform similarity analysis based on the input data within different pre-similar datasets, perform dual matching of high-frequency field overlap and intent relevance on the input data within the pre-similar datasets, analyze and build similar datasets, and build a group data feature profile of similar datasets based on different data categories;
[0046] according to Figure 3 Furthermore, it needs to be explained in detail that the multi-category data dynamic clustering module includes a multi-dimensional comparison and discrimination module for similar datasets and a pre-similar data feature similarity analysis and clustering sub-module. The multi-dimensional comparison and discrimination module for similar datasets acquires pre-processed data to be approved, extracts arbitrary keywords from the text data within the input data to be approved, marks key data features, performs pre-comparison based on the marked key data features of different input data, analyzes the overlap of key data features based on different input data, classifies input data with key data feature overlap exceeding a set threshold, and marks them as pre-similar input datasets, screens high-frequency words with pre-similar input datasets, performs intent relevance judgment based on the screened high-frequency words, and screens the intent relevance between each unclassified input data and each pre-similar input dataset. When the intent relevance between the unclassified input data and each pre-similar input dataset is greater than a set threshold, the input data is classified into the pre-similar input dataset.
[0047] To further explain, the pre-class data feature similarity analysis clustering submodule obtains each pre-class dataset, analyzes the similarity of the input data within each pre-class dataset, and sets the similarity score of the input data within each pre-class dataset as follows: The input data within each pre-class dataset is matched against the high-frequency words in the pre-class dataset using field overlap matching. The overlap rate between the input data within each pre-class dataset and the high-frequency words is set to a certain percentage. The process involves filtering out data fields within each pre-class dataset that overlap with high-frequency word fields. Then, it performs intent relevance matching between the remaining fields of the input data within each pre-class dataset and the intents related to each high-frequency word in the pre-class dataset. The intent relevance between the remaining fields of the input data within each pre-class dataset and the intents related to each high-frequency word in the pre-class dataset is set to [value missing]. A multi-factor weighted model was used to perform similarity analysis between the input data within different pre-class datasets and the dataset of the same category. The weighting of the high-frequency words in the input data and the pre-class input datasets was set to account for overlap in the matching of fields. The input data is matched with high-frequency words in a pre-classified input dataset based on intent association, with the weighting ratio being... Set the similarity score threshold for input data within each pre-class dataset to be [value]. According to the formula:
[0048]
[0049] The similarity score of each input data set within a pre-class dataset is calculated to determine the similarity between the input data and the pre-class input dataset. If the above formula is satisfied, it means that the input data belongs to the same category within the pre-class input dataset. When the similarity score of each input data belongs to the same category... If the above formula is not met, the input data is removed and reclassified until the similarity score between the input data and the corresponding pre-classified dataset meets the above threshold. Each input data within each pre-classified dataset is screened, and if the similarity score between the input data and the corresponding pre-classified dataset meets the above threshold, it is re-labeled. The pre-classified dataset corresponding to this type of input data is labeled as a class dataset. Based on the category of each class dataset, a group text data feature profile of the class dataset is constructed.
[0050] The multi-data feature association deviation analysis module is used to associate features of image data and special data of each input data in the same category dataset, analyze and screen out based on the feature differences of image data and special data of different input data, and measure the degree of deviation of input data from the group mean in semantic information features;
[0051] according to Figure 4 Furthermore, it needs to be explained in detail that the multi-data feature association deviation analysis module includes a data feature dynamic association identification screening module and an association data deviation value quantification analysis submodule. The data feature dynamic association identification screening module is used to acquire image data and special data of each input data in the same dataset, pre-identify text regions and image regions in the image data, extract text information within the text regions, convert it into text semantic vectors through a natural language processing model, extract the image part of the image region for content recognition, and determine whether the recognized content of the image region matches the text information of the text region. If they match, extract the group text data feature profile within the same dataset, compare the text semantic information of the text region with the group text data feature profile within the same dataset, analyze the correlation between the image data of each input data and the group text data feature profile within the same dataset, and filter out data in the same dataset where the correlation between the image data of each input data and the group text data feature profile within the same dataset is less than a set threshold.
[0052] Obtain the special data of each input data in the same dataset, extract the statistical features, numerical distribution features and syntactic structure features of the special data respectively, compare the features of the special data of the input data in each dataset, and filter out the input data corresponding to the obvious abnormal special data features.
[0053] The input data corresponding to the screening of image data and the feature profiles of the group text data within the same dataset that have a correlation degree of less than a set threshold and the presence of obvious abnormal special data features are statistically analyzed. The input data is sent to the manual approval and optimization platform for review. The same dataset after the screening operation is obtained, and the same dataset is marked three times to become the verification dataset.
[0054] To further explain, the correlation data deviation quantification analysis submodule acquires each input data within the same verification dataset, obtains the corresponding group text data feature profile of the same verification dataset, extracts each core feature within the corresponding group text data feature profile, extracts the semantic information of each core feature, compares the text data semantic information features of the input data and the text semantic information features of the text region in the image data with any core feature using single feature deviation, and the deviation judgment rule is uploaded by the user. The average deviation of the text data semantic information features of each input data and the text semantic information features of the text region in the image data with the semantic information features of each core feature is analyzed separately. The average deviation of the text data semantic information features of each input data and the text semantic information features of the text region in the image data with the semantic information features of each core feature is combined to calculate the comprehensive deviation score of the input data, and a reasonable range for the comprehensive deviation is set. Set the overall deviation score of a certain input data as Set the deviation quantization value for different input data. According to the formula:
[0055]
[0056] Obtain the deviation quantization values of different input data within the same dataset for verification, and divide the different input data within the same dataset for verification based on the different deviation quantization values.
[0057] The multi-type approval rule matching module matches the corresponding approval rules based on the data of the same category, performs intelligent verification on the data of the same category in sequence according to the matched approval rules, performs random secondary approval feasibility verification based on the data approval results, and provides feedback on abnormal approval data.
[0058] according to Figure 5 Furthermore, it needs to be explained in detail that the multi-type approval rule matching module includes a multi-dataset approval rule matching sub-module, an approval strategy feasibility call verification sub-module, and an abnormal approval reporting feedback sub-module. The multi-dataset approval rule matching sub-module is used to obtain the feature profile of the group text data of the same type of dataset for verification, obtain the uploaded approval rule library data, match different approval rules according to the feature profile of the group text data, and approve the input data within each verification of the same type of dataset one by one according to different approval rules, and generate approval results. If the approval result of a certain input data fails, the final approval opinion is generated according to the corresponding approval rule. The approval opinion includes at least one specific error type and misalignment location information, and is sent to the multi-path approval error correction deduction and verification module.
[0059] The feasibility verification submodule for the approval strategy calls upon the user to obtain input data with deviation quantization values of 1 and 0 from the same verification dataset. It then extracts input data with deviation quantization values of 1 and 0 from each verification dataset for secondary approval verification. The amount of input data with deviation quantization value of 1 is set to be... The input data volume is 1, with a deviation quantization value of 1. , The system screens the first and second approval results of input data with deviation values of 1 and 0 in each sampled dataset of the same type of data. It verifies whether the first and second approval results of the same input data are consistent. If the approval results are consistent, it marks that the verification strategy for matching the same type of data is effective. If the approval results are inconsistent, it screens the inconsistent data and marks them as abnormal. The abnormal approval reporting feedback submodule sends the abnormally marked input data to the manual approval optimization platform for manual approval.
[0060] The multi-path approval error correction derivation and verification module locates erroneous data based on different data approval correction opinions, provides error correction prompts based on the data error type, performs confidence analysis on each of the multiple error correction prompts, and uploads a backup after user feedback confirmation on the error correction prompts after confidence verification.
[0061] according to Figure 6 Furthermore, it needs to be explained in detail that the multi-path approval error correction derivation and verification module includes a multi-data approval correction and error correction data marking sub-module, a correction value confidence analysis verification sub-module, and a correction result manual feedback optimization confirmation sub-module. The multi-data approval correction and error correction data marking sub-module obtains the final approval opinions based on several input data within each verification similar dataset. Based on the final approval opinions of different input data, including error type and misalignment location information, it determines the error type of the input data. Error types include format errors, logical errors, partial data errors, and partial information missing. For format errors and logical errors, it corrects them according to the correction database uploaded by the manual approval optimization platform. It summarizes the corrected input data and sends the corrected input data to the corresponding approval rules in the approval rule base for secondary approval. When the approval is passed, it replaces the original input data within the verification similar dataset with the corrected input data to form a new verification similar dataset. For input data with partial data errors and partial information missing, it obtains the historical sending data of the user who sent the input data and finds the K historical sending data records that are most similar to the target input data, using the mode of their corresponding fields as candidate missing data.
[0062] The modified value confidence analysis verification submodule obtains candidate missing data from the target input data, retrieves the historical data path of the candidate missing data, analyzes the number of times it appears in the historical data, and sets the number of times it appears in the historical data as [value missing]. The baseline confidence level for its candidate missing data is set to 1. , = The candidate missing data is mapped to the context of the target input data, and the matching degree between the candidate missing data and the target input data context is analyzed. The number of context fields in the target input data that are logically related to the candidate missing data is set to 1. The contextual fit between the candidate missing data and the target input data is set to 1. According to the formula:
[0063]
[0064] In the above formula, For candidate missing data values, After filling in the missing candidate data in the target input data, the new input data is formed with the first... The values of the context fields This is a consistency indicator function used to determine the collected candidate missing data values. With any context field Are they logically similar when collecting candidate missing data values? and If the candidate missing data values are similar and logically consistent, the function output will be 1. and If they are different and logically contradictory, the function output will be 0;
[0065] Analyze the final confidence level of the candidate missing data, and set the final confidence level of the candidate missing data as . According to the formula:
[0066]
[0067] in, and These are the base confidence level of manually set candidate missing data and the configuration weight coefficients for the contextual fit between candidate missing data and target input data, respectively. ,when If the value exceeds a set threshold, the candidate missing data is substituted into the corresponding input data to replace the erroneous fields and fill in the missing information. It is then marked as pre-corrected data. The pre-corrected data is sent to the data sending user for confirmation through the manual feedback optimization confirmation submodule. After confirmation, the input data in the corresponding similar verification dataset is replaced to form a new similar verification dataset. The new similar verification dataset is then sent to the manual approval optimization platform for backup.
[0068] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A large-scale master data automatic approval error correction system based on an agent, characterized in that: The system comprises an approval data access verification module, a multi-category data dynamic clustering module, a multi-data feature correlation deviation analysis module, a multi-type approval rule matching corresponding module, a multi-path approval error correction deduction re-inspection module and an artificial approval optimization platform, and the approval data access verification module, the multi-category data dynamic clustering module, the multi-data feature correlation deviation analysis module, the multi-type approval rule matching corresponding module, the multi-path approval error correction deduction re-inspection module and the artificial approval optimization platform are sequentially connected; The approval data access verification module is used for preprocessing the approval data input by the approval data source, performing field-based preliminary compliance verification on the data source, and performing standardized verification processing on the multi-field format of the input data; The multi-category data dynamic clustering module is used for multi-dimensional comparison of key fields in the text of different to-be-approved data, constructing a dynamic clustering model, dynamically clustering different category data, automatically dividing different input data into specific data categories, constructing a pre-homogeneous data set, performing similarity analysis on the internal input data of different pre-homogeneous data sets, performing double matching of high-frequency field overlap and intention correlation on the internal input data of the pre-homogeneous data set, analyzing and constructing a homogeneous data set, and constructing the group data feature portrait of the homogeneous data set according to different data categories; The multi-data feature correlation deviation analysis module is used for feature correlation of image data and special data of each input data in the same category data set, and analyzes and screens out the feature differences of the image data and special data of different input data to measure the degree of deviation of the input data from the group mean value in semantic information features; The multi-type approval rule matching corresponding module matches the corresponding approval rules according to the same category data, sequentially intelligently verifies the same category data according to the matched approval rules, performs random secondary approval feasibility verification according to the data approval result, and feeds back the abnormal approval data; The multi-path approval error correction deduction re-inspection module locates error data according to different data approval correction opinions, provides error correction prompts according to the data error types, performs one-by-one confidence analysis according to multiple error correction prompts, and uploads backups after user feedback confirmation of the error correction prompts after confidence re-inspection.
2. The agent-based large-scale master data automatic approval and correction system according to claim 1, characterized in that The approval data access verification module comprises a multi-data source data preprocessing submodule and a data format standardization verification submodule. The multi-data source data preprocessing submodule obtains original data sent by different data sources, cleans and preprocesses the original data, verifies input data format, data sending time and data sending user, eliminates data missing and data format errors with obvious abnormalities, eliminates data with abnormal sending time and unclear sending user, and the data format standardization verification submodule obtains preprocessed data, divides different data according to different format fields, and screens different fields according to different standardization format rules. The standardization format rules are uploaded by humans, the non-standard rate of each input data field format is screened, when the non-standard rate of a field format is greater than a set threshold, the average proportion of the total non-standard rate of the input data is obtained, when the average proportion of the non-standard rate of the input data is greater than a set threshold, the input data is fed back to the sending user for correction.
3. The agent-based large-scale master data automatic approval and correction system according to claim 1, characterized in that The multi-category data dynamic clustering module comprises a same data set multi-dimensional comparison and discrimination module and a pre-same data feature similarity analysis and clustering submodule. The same data set multi-dimensional comparison and discrimination module obtains preprocessed data to be approved, pre-extracts any key words from the text data in the data to be approved, marks key data features, pre-compares different input data according to the marked key data features, analyzes the key data feature overlap degree of different input data, classifies input data with a key data feature overlap degree higher than a set threshold, marks the input data as a pre-same input data set, screens high-frequency words of the pre-same input data set, makes an intention correlation judgment according to the screened high-frequency words, screens the intention correlation degree of each unclassified input data and each pre-same input data set, and classifies the input data into the pre-same input data set when the intention correlation degree of the unclassified input data and each pre-same input data set is greater than a set threshold.
4. The agent-based large-scale master data automatic approval and correction system according to claim 3, characterized in that The pre-homogeneous data feature similarity analysis clustering submodule obtains each pre-homogeneous data set, analyzes the homogeneity similarity of the input data in each pre-homogeneous data set, sets the homogeneity similarity score of the input data in each pre-homogeneous data set as , performs field overlap matching on each pre-homogeneous data set and the high-frequency words of the pre-homogeneous input data set, sets the field matching overlap rate of each pre-homogeneous data set and the high-frequency words as , screens out the data fields of each pre-homogeneous data set that have been matched and overlapped with the high-frequency words field, performs intent association degree matching on the remaining fields of each pre-homogeneous data set and each high-frequency word related intent of the pre-homogeneous input data set, and sets the intent matching association degree of the remaining fields of each pre-homogeneous data set and each high-frequency word related intent of the pre-homogeneous input data set as , a multi-factor weighting model is used to analyze the similarity of the input data in different pre-homogeneous data sets and the data set of the category, sets the field overlap matching weight proportion of the input data and the high-frequency words of the pre-homogeneous input data set as , the intent association matching weight proportion of the input data and the high-frequency words of the pre-homogeneous input data set is , sets the homogeneity similarity score threshold of each pre-homogeneous data set as , according to the formula: The similarity score of each input data set within a pre-class dataset is calculated to determine the similarity between the input data and the pre-class input dataset. If the above formula is satisfied, it means that the input data belongs to the same category within the pre-class input dataset. When the similarity score of each input data belongs to the same category... If the above formula is not met, the input data is removed and reclassified until the similarity score between the input data and the corresponding pre-classified dataset meets the above threshold. Each input data within each pre-classified dataset is screened, and if the similarity score between the input data and the corresponding pre-classified dataset meets the above threshold, it is re-labeled. The pre-classified dataset corresponding to this type of input data is labeled as a class dataset. Based on the category of each class dataset, a group text data feature profile of the class dataset is constructed.
5. The agent-based large-scale master data automatic approval and correction system according to claim 1, characterized in that The multi-data feature correlation deviation analysis module comprises a data feature dynamic correlation identification screening submodule and a correlation data deviation value quantization analysis submodule. The data feature dynamic correlation identification screening submodule is configured to acquire image data and special data of each input data in the same data set, pre-identify a text region and an image region in the image data, extract text information in the text region, convert the text information into a text semantic vector through a natural language processing model, extract an image part of the image region for content identification, determine whether the identified content of the image region is information-correlation-matched with the text information of the text region, if matched, extract a group text data feature portrait in the same data set, compare the text semantic information of the text region with the group text data feature portrait in the same data set, analyze the correlation degree of the image data of each input data with the group text data feature portrait in the same data set, and screen out data whose correlation degree is less than a set threshold value; acquire special data of each input data in the same data set, respectively extract statistical features, numerical distribution features and syntax structure features of the special data, compare the special data features of the input data in each same data set, and screen out input data with obviously abnormal special data features; respectively count the input data with the correlation degree of the image data with the group text data feature portrait in the same data set being less than the set threshold value and the input data with obviously abnormal special data features, send the input data to an artificial approval optimization platform for review, acquire the same data set after the screening operation, mark the same data set three times, and regard the same data set as a verification same data set.
6. The agent-based large-scale master data automatic approval and correction system according to claim 5, characterized in that The correlation data deviation value quantitative analysis submodule obtains each input data in the verification same kind data set, obtains the corresponding group text data feature image of the verification same kind data set, extracts each core feature in the corresponding group text data feature image, extracts the semantic information of each core feature, compares the text data semantic information feature of the input data and the text semantic information feature of the picture number drama text area with any core feature in single feature deviation degree, and the deviation degree determination rule is uploaded by thinking. The average deviation degree of the text data semantic information feature of each input data and the text semantic information feature of the picture number drama text area and the semantic information feature of each core feature is analyzed respectively, the average deviation degree of the text data semantic information feature of each input data and the text semantic information feature of the picture number drama text area and the semantic information feature of each core feature is combined, the comprehensive deviation degree score of the input data is calculated, the reasonable interval of the comprehensive deviation degree is set as , the comprehensive deviation degree score of a certain input data is set as , the deviation quantitative value of different input data is set as , according to the formula: respectively acquire deviation quantization values of different input data in the verification same data set, and divide the different input data in the same verification same data set according to different deviation quantization values.
7. The agent-based large-scale master data automatic approval and correction system according to claim 1, characterized in that The multi-type approval rule matching corresponding module comprises a multi-data set approval rule corresponding submodule, an approval strategy feasibility calling verification submodule and an abnormal approval reporting feedback submodule. The multi-data set approval rule corresponding submodule is configured to acquire a group text data feature portrait of the verification same data set, acquire uploaded approval rule library data, match different approval rules according to the group text data feature portrait, perform approval on each input data in the verification same data set one by one according to different approval rules, generate an approval result, generate a final approval opinion according to the corresponding approval rule if the approval result of the input data does not pass, and send the approval opinion to the multi-path approval error correction derivation re-inspection module. The approval strategy feasibility calling verification submodule user respectively acquires the input data of the deviation quantization value 1 and the deviation quantization value 0 in the same check same kind data set, respectively extracts the input data of the deviation quantization value 1 and the deviation quantization value 0 in each check same kind data set for secondary approval verification, sets the extracted input data amount of the deviation quantization value 1 as , the extracted input data amount of the deviation quantization value 1 as , , screens the first approval result and the second approval result of the input data of the deviation quantization value 1 and the deviation quantization value 0 in each check same kind data set, verifies whether the first approval result and the second approval result of the same input data are consistent, when the approval result is consistent, marks the verification strategy of the check same kind data matching as effective, when the approval result is inconsistent, screens the data inconsistent in the approval for abnormal marking, and sends the input data of the abnormal marking to the artificial approval optimization platform for artificial approval through the abnormal approval feedback submodule.
8. The agent-based large-scale master data automatic approval and correction system according to claim 1, characterized in that The multi-path approval error correction derivation rechecking module comprises a multi-data approval correction error correction data marking submodule, a correction value confidence analysis rechecking submodule, and a correction result manual feedback optimization confirmation submodule. The multi-data approval correction error correction data marking submodule obtains final approval opinions according to a plurality of input data in each check similar data set, and the final approval opinions of different input data include error types and error location information. The error types of the input data are judged, and the error types include format errors, logic errors, partial data errors, and partial information missing. The format errors and the logic errors are corrected according to a correction database uploaded by a manual approval optimization platform. The corrected input data is summarized and sent to corresponding approval rules in an approval rule library for secondary approval. When the approval is passed, the corrected input data is replaced with original input data in the check similar data set to form a new check similar data set. For the input data with partial data errors and partial information missing, historical sending data of a sending user of the input data is obtained, K historical sending data records most similar to the target input data are found, and the mode of corresponding fields is used as candidate missing data. The correction value confidence analysis rechecking sub-module acquires candidate missing data of target input data, acquires historical data path of the candidate missing data, analyzes occurrence frequency of the candidate missing data in historical sending data, sets the occurrence frequency of the candidate missing data in historical sending data as , sets basic confidence of the candidate missing data as , = , corresponds the candidate missing data with context of target input data, analyzes fitting degree of the candidate missing data with context of target input data, sets context field number of target input data which is logically related to the candidate missing data as , sets fitting degree of the candidate missing data with context of target input data as , according to formula: In the above formula, is a candidate missing data value, is a target input data after filling the candidate missing data, is the value of the first context field in the new input data, is a consistency indication function, which judges whether the collected candidate missing data value is logically similar to any context field, when the collected candidate missing data value is similar to and logically consistent, the function outputs a result of 1, and when the collected candidate missing data value is different from and logically contradictory, the function outputs a result of 0. analyzing the final confidence of the candidate missing data, setting the final confidence of the candidate missing data as , according to the formula: wherein, and are respectively the basic confidence of the candidate missing data and the configuration weight coefficient of the candidate missing data and the target input data context, when is greater than a set threshold, the candidate missing data is substituted into the corresponding input data to replace the error field and fill in the missing information, marked as pre-corrected data, the pre-corrected data is sent to the data sending user for confirmation by the correction result manual feedback optimization confirmation submodule, after confirmation, the corresponding input data in the new check similar data set is replaced, a new check similar data set is formed, and the new check similar data set is sent to the artificial approval optimization platform for backup.
Citation Information
Patent Citations
Production operation management system
CN118410917A
Application form intelligent auditing and approving method and system and medium
CN120338349A