Document information entry method and system based on big data processing

Through big data processing technology, compliance data objects in documents are identified, task scheduling weight parameters are generated, and multi-dimensional compliance verification and risk scanning are carried out in combination with knowledge graphs and dynamic rule bases. This solves the problems of delayed rule updates and unreasonable audits in document information entry, and achieves efficient and accurate document information entry and auditing.

CN119990719BActive Publication Date: 2025-09-19SHENZHEN TEWEI KECHUANG INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510477952.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-09-19
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Existing technologies make it difficult to fully utilize big data mining technology to conduct in-depth risk analysis during the document information entry process, resulting in false positives or omissions, delayed rule updates, and unreasonable allocation of audit tasks, affecting efficiency and accuracy.

Method used

The document information entry method based on big data processing identifies compliance data objects, generates task scheduling weight parameters, combines knowledge graphs and dynamic rule knowledge bases, conducts multi-dimensional compliance verification and risk feature scanning, adopts intelligent task routing and multi-level verification, generates credibility scores, and optimizes the audit path.

Benefits of technology

It achieves efficient and accurate document information entry, improves the flexibility and timeliness of compliance checks, optimizes the audit path, reduces manual intervention, ensures the rigor and reliability of data processing, and supports automated risk warning and task allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990719B_ABST
    Figure CN119990719B_ABST
Patent Text Reader

Abstract

The present invention discloses a document information entry method and system based on big data processing, which belongs to the field of document information processing technology. The present invention combines big data mining and knowledge graphs to conduct in-depth risk feature screening of document data, can accurately identify potential sensitive factors, and analyze risk categories and impact ranges based on association rules, realize automatic early warning and real-time risk notification, adopt a dynamic rule knowledge base, support multi-dimensional compliance verification, can update business constraints in real time, ensure the flexibility and adaptability of rules, improve the accuracy and timeliness of compliance inspections, through intelligent verification processes, use the review rule set trained by big data, perform pattern matching and multi-level verification on compliance data objects, ensure the rigor and reliability of data auditing, generate task scheduling weight parameters through comprehensive quantitative calculation of business complexity, risk coefficient and timeliness priority, optimize document audit path, and improve audit efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document information processing, and in particular to a document information entry method and system based on big data processing. Background Art

[0002] In the context of modern data management and business process automation, various enterprises, financial institutions, and government departments need to process large amounts of document information, such as invoices, contracts, expense reimbursements, and bank statements. These documents often come from a wide range of sources and in various formats, including both structured and unstructured or semi-structured data. With increasing corporate compliance and regulatory requirements, document data entry requires rigorous risk screening and compliance verification. Currently, existing technologies still have some limitations, such as:

[0003] Traditional risk control methods often rely on preset rules and cannot fully utilize big data mining technology for in-depth risk analysis, which is prone to false positives or omissions. Some existing systems use fixed rule sets for compliance checks, which are difficult to dynamically adapt to changing business rules, resulting in delayed rule updates or incomplete verification. The allocation of audit tasks usually relies on manual judgment or simple fixed rules, and fails to combine business complexity, risk level and audit priority for intelligent scheduling, affecting audit efficiency and accuracy. Summary of the Invention

[0004] The purpose of the present invention is to provide a document information entry method and system based on big data processing to solve the problems raised in the above background technology.

[0005] To achieve the above-mentioned purpose, the present invention provides the following technical solution: a document information entry method based on big data processing, comprising:

[0006] Collecting original document data and identifying compliant data objects in the original document data, generating task scheduling weight parameters for the compliant data objects, wherein the process of determining the task scheduling weight parameters includes: retrieving a correlation value based on a data type and a key data field of the compliant data object, obtaining a first complexity coefficient and a second complexity coefficient based on the correlation value, and generating the task scheduling weight parameter in combination with a risk coefficient and a timeliness priority quantified value;

[0007] Based on the task scheduling weight parameters, each compliant data object is matched to the optimal audit path. After the compliant data object is audited through the optimal audit path, the final processing result is output and archived to the business database.

[0008] Furthermore, it includes: extracting original document data stored in diverse media carriers, aligning the original document data to a template framework and generating a structured document data object, wherein the diverse media carriers include database clusters, optical character recognition files, electronic communication attachment documents and IoT device record streams.

[0009] Furthermore, it includes: scanning the risk characteristics of document data objects based on big data mining and combining with knowledge graphs to identify sensitive elements in the preset risk characteristic library;

[0010] Analyze the characteristics of document data objects based on big data mining technology, extract key information elements from the document, compare the extracted key information elements with the preset risk feature library, and identify possible sensitive elements;

[0011] If sensitive elements are detected, the process will be blocked, and the risk category and impact scope will be determined based on association rule analysis. Real-time warnings will be issued and multi-channel risk notifications will be pushed to the monitoring terminal.

[0012] Furthermore, it includes: performing multi-dimensional compliance verification on document data objects. If any dimension verification fails, the process is blocked and a visual diagnostic report is generated to assist operators in locating the root cause of the anomaly. If the verification succeeds, the compliant data object is output;

[0013] Build a dynamic rule knowledge base, extract business constraints for each verification dimension from the dynamic rule knowledge base, and identify rule version status, which includes existing rules and newly added rules;

[0014] If it is an existing rule, the document data object is pattern matched with the business constraints of the corresponding dimension. If the business constraints are met, the rule is considered compliant. If it is a new rule, the compatibility of the logical expression of the business constraint and the input parameters is checked. If there is a compatibility conflict, the conflicting business constraint is frozen.

[0015] When a business constraint is detected as a newly added rule, check whether the business constraint covers all data dimensions. If any dimension is missing, suspend the business constraint from taking effect.

[0016] Further, it includes: implementing multi-level verification of compliant data objects based on a review rule set trained on big data, identifying abnormal patterns and generating credibility scores;

[0017] Based on a review rule set trained on big data, pattern matching is performed on compliant data objects to extract abnormal features. Combined with a multi-level verification mechanism, rule constraint verification is performed on compliant data objects, and a credibility score is calculated for compliant data objects. The pass / fail criteria are then set based on preset score thresholds.

[0018] If the verification fails, the process will be blocked and an abnormal pattern analysis report will be generated. If the verification passes, the compliant data object will be output to the next stage.

[0019] Furthermore, it includes: intelligent task routing for compliant data objects that have passed verification, matching each compliant data object to the optimal review path based on multiple criteria such as business complexity, risk factor, and timeliness priority;

[0020] Quantify the business complexity, risk factor, and timeliness priority of compliance data objects, generate task scheduling weight parameters, and match compliance data objects to the optimal review path based on the task scheduling weight parameters;

[0021] Generate audit tasks based on compliance data objects, assign audit tasks to adapted audit terminals, and after the audit is completed, output the final processing results and archive them in the business database.

[0022] Furthermore, the quantitative calculation of the business complexity, risk factor, and timeliness priority of the compliant data objects to generate task scheduling weight parameters includes:

[0023] Extracting the data types contained in the compliant data object and the key data fields contained in each data type;

[0024] Retrieve the preset correlation value between the key data field contained in each data type and other key data fields of the same data type within the data type;

[0025] Obtaining a first complexity coefficient corresponding to each data type using a preset correlation value between a key data field contained in each data type and other key data fields of the same data type within the data type to which the key data field belongs;

[0026] The first complexity coefficient is obtained by the following formula:

[0027]

[0028] Among them, R 01 Indicates the first complex coefficient corresponding to each data type; n indicates the number of key data fields; L bi Indicates the standard deviation of the correlation between the i-th key data field and other key data fields of the same data type within the data type in which it is located; L maxi Indicates the maximum value of the correlation between the i-th key data field and other key data fields of the same data type within the data type in which it is located; L miniIndicates the minimum value of the correlation between the i-th key data field and other key data fields of the same data type within the data type in which it is located; x i Indicates the number of maximum correlation values ​​corresponding to the i-th key data field; y i Indicates the number of minimum correlation values ​​corresponding to the i-th key data field;

[0029] Retrieve the preset correlation values ​​between the key data fields contained in each data type and other data types;

[0030] Obtaining a second complexity coefficient corresponding to each data type by utilizing the first complexity coefficient corresponding to each data type in combination with a preset correlation value between the key data field contained in each data type and other data types;

[0031] The second complexity coefficient is obtained by the following formula:

[0032]

[0033] Among them, R 02 Indicates the second complex coefficient corresponding to each data type; R 01 Indicates the first complex coefficient corresponding to each data type; n indicates the number of key data fields; L bgi Indicates the standard deviation of the preset correlation between the i-th key data field and other data types; L gmaxi Indicates the maximum value of the correlation between the i-th key data field and other data types; L gmini Indicates the minimum value of the correlation between the i-th key data field and other data types; A i Indicates the number of maximum values ​​of the correlation between the i-th key data field and other data types; B i Indicates the number of minimum values ​​of the correlation between the i-th key data field and other data types; Z 01i Indicates the number of other data types included when the correlation value between the i-th key data field and other data types is the maximum; Z 02i Indicates the number of other data types included when the correlation value between the i-th key data field and other data types is the maximum;

[0034] The task scheduling weight parameter is generated by utilizing the risk coefficient corresponding to the compliant data object and the quantitative value of the timeliness priority in combination with the second complex coefficient.

[0035] Furthermore, the risk coefficient and the quantified value of the timeliness priority corresponding to the compliant data object are combined with the second complex coefficient to generate a task scheduling weight parameter, including:

[0036] Identifying risk events contained in the compliance data objects;

[0037] Performing a risk probability analysis on the risk events to obtain the corresponding probability of occurrence of each risk event;

[0038] The risk factor corresponding to each compliant data object is obtained by performing a weighted average using the probability of occurrence of the risk event and the weight parameter corresponding to the risk event;

[0039] Identify key time nodes in the processing of compliant data objects and obtain the time difference between each key time node and the current time;

[0040] Identify the incidence of risk events corresponding to the key time nodes;

[0041] The quantitative value of the timeliness priority is obtained by using the occurrence rate of the risk events corresponding to the key time nodes and the time difference between each key time node and the current time;

[0042] The quantified value of the timeliness priority is obtained by the following formula:

[0043]

[0044] Among them, F represents the quantitative value of time priority; m represents the number of key time nodes; P i represents the occurrence rate of risk events corresponding to the i-th key time node; T i Indicates the time difference between the i-th key time node and the current time; T c Indicates the preset time difference reference value;

[0045] Retrieve the second complex coefficient;

[0046] The risk coefficient and the quantified value of the timeliness priority corresponding to the compliant data object are combined with the second complex coefficient to generate an initial task scheduling weight parameter; wherein the initial task scheduling weight parameter is obtained by the following formula:

[0047]

[0048] Among them, W ctask represents the initial task scheduling weight parameter; k represents the number of data types contained in the compliant data object; R 02irepresents the second complexity coefficient corresponding to the i-th data type; F represents the quantitative value of the timeliness priority; U represents the risk coefficient corresponding to the compliance data object;

[0049] The initial task scheduling weight parameter corresponding to each compliant data object is standardized to generate the task scheduling weight parameter corresponding to each compliant data object.

[0050] Furthermore, a document information entry system based on big data processing is applied to the above-mentioned document information entry method based on big data processing, including:

[0051] Heterogeneous data source parsing module, used for:

[0052] Extract original document data stored in various media carriers, parse document data in different formats, and generate standardized data input streams;

[0053] Intelligent data mapping module for:

[0054] Receive the parsed original document data, align the data according to the template framework, and generate a structured document data object;

[0055] Risk signature screening module for:

[0056] Scan document data objects for risk characteristics, analyze their characteristics, extract key information elements, and identify potential sensitive elements based on these key information elements;

[0057] Multi-dimensional compliance verification module, used for:

[0058] Extract business constraints for each verification dimension from the dynamic rule knowledge base and identify the rule version status, and perform multi-dimensional compliance verification on document data objects based on the business constraints.

[0059] Furthermore, the document information entry system based on big data processing also includes:

[0060] Intelligent verification process module for:

[0061] Based on the review rule set trained on big data, multi-level verification is carried out on compliant data objects, abnormal features are extracted, rule constraints are verified on compliant data objects, and the credibility score of compliant data objects is calculated;

[0062] Audit task scheduling module, used to:

[0063] Quantify the business complexity, risk factor, and timeliness priority of compliance data objects, generate task scheduling weight parameters, and match compliance data objects to the optimal review path based on the task scheduling weight parameters;

[0064] Assign audit tasks to the audit terminal, and archive compliant data objects to the business database after the audit is completed.

[0065] Compared with the prior art, the present invention has the following beneficial effects:

[0066] 1. This invention uses heterogeneous data source parsing technology to efficiently extract original document data stored in databases, OCR files, electronic communication accessories, and IoT device record streams, enabling automatic parsing and formatting of multi-source data. Through intelligent data mapping, the parsed unstructured or semi-structured document data is converted into standardized structured data objects, improving the consistency and accuracy of data processing and reducing manual intervention.

[0067] 2. The present invention combines big data mining and knowledge graphs to conduct in-depth risk feature screening of document data, accurately identify potential sensitive factors, and analyze risk categories and impact scope based on association rules to achieve automatic early warning and real-time risk notification. It adopts a dynamic rule knowledge base, supports multi-dimensional compliance verification, and can update business constraints in real time to ensure the flexibility and adaptability of rules, improve the accuracy and timeliness of compliance checks, and through an intelligent verification process, use the review rule set trained by big data to perform pattern matching and multi-level verification on compliant data objects, and generate credibility scores to ensure the rigor and reliability of data audits.

[0068] 3. The present invention generates task scheduling weight parameters through comprehensive quantitative calculation of business complexity, risk coefficient and timeliness priority, optimizes the document review path, improves review efficiency, adopts a dynamic load balancing algorithm, realizes intelligent task allocation, ensures the optimal utilization of review resources, reduces manual review pressure, improves overall business processing capabilities, supports automatic archiving and recording of review tasks, ensures the integrity and traceability of data tracking, and meets regulatory and audit requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 Schematic diagram of the document information entry method of the present invention;

[0070] Figure 2 This is a schematic diagram of the document information entry process of the present invention;

[0071] Figure 3 This is a schematic diagram of the document information entry system module of the present invention. DETAILED DESCRIPTION

[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0073] See also Figure 1-Figure 2 , the present invention provides the following technical solutions:

[0074] The document information entry method based on big data processing includes:

[0075] Collecting original document data and identifying compliant data objects in the original document data, generating task scheduling weight parameters for the compliant data objects, wherein the process of determining the task scheduling weight parameters includes: retrieving a correlation value based on a data type and a key data field of the compliant data object, obtaining a first complexity coefficient and a second complexity coefficient based on the correlation value, and generating the task scheduling weight parameter in combination with a risk coefficient and a timeliness priority quantified value;

[0076] Based on the task scheduling weight parameters, each compliant data object is matched to the optimal audit path. After the compliant data object is audited through the optimal audit path, the final processing result is output and archived to the business database.

[0077] The specific steps include:

[0078] Heterogeneous data source parsing phase: Extracting original document data stored in diverse media carriers, including database clusters, optical character recognition files, electronic communication attachments, and IoT device record streams;

[0079] Intelligent data mapping stage: align the parsed original document data to the template framework and generate structured document data objects;

[0080] Risk feature screening stage: Based on big data mining and combined with knowledge graphs, the document data objects are scanned for risk features to identify sensitive elements in the preset risk feature library;

[0081] Multi-dimensional compliance verification stage: Perform multi-dimensional compliance verification on document data objects. If any dimension verification fails, the process is blocked and a visual diagnostic report is generated to assist operators in locating the root cause of the anomaly. If the verification succeeds, the compliant data object is output.

[0082] Intelligent verification process stage: Based on the review rule set trained on big data, multi-level verification is carried out on compliant data objects to identify abnormal patterns and generate credibility scores;

[0083] Audit task scheduling stage: Intelligent task routing is performed on compliant data objects that have passed verification, and each compliant data object is matched to the optimal audit path based on multiple criteria such as business complexity, risk factor, and time priority.

[0084] In the above embodiment, through intelligent data mapping, the parsed data is automatically aligned to the preset template framework, reducing the workload of manual adjustment and the input error rate. By utilizing big data mining and knowledge graphs, it is possible to efficiently detect risk features in document data, accurately identify sensitive information, and improve risk management capabilities. Through a dynamic rule knowledge base, the system can adapt to changes in business rules, achieve accurate rule constraint checks, and improve the accuracy of compliance judgments. By using a review rule set based on big data training, it can automatically identify abnormal patterns, calculate credibility scores, and ensure data quality and consistency. Audit task scheduling uses an intelligent routing algorithm, combined with business complexity, risk factors, and timeliness priorities to ensure efficient and reasonable allocation of audit tasks and optimize business processes.

[0085] The risk profile screening phase also includes:

[0086] Analyze the characteristics of document data objects based on big data mining technology, extract key information elements from the document, compare the extracted key information elements with the preset risk feature library, and identify possible sensitive elements;

[0087] If sensitive elements are detected, the process will be blocked, and the risk category and impact scope will be determined based on association rule analysis. Real-time warnings will be issued and multi-channel risk notifications will be pushed to the monitoring terminal.

[0088] In the above embodiment, through big data mining and knowledge graph technology, the risk characteristics of the document data object are scanned to identify potential risk factors and improve the security of document processing. The use of knowledge graph and risk characteristic library can efficiently identify abnormal information in the document, such as financial fraud, data leakage and other risks. Once sensitive factors are detected, the system will immediately trigger the process blocking and push risk notifications to the risk control terminal to ensure timely intervention. Through big data mining technology, the risk category and impact range can be analyzed to improve the accuracy of risk control.

[0089] The multi-dimensional compliance verification stage further includes:

[0090] Build a dynamic rule knowledge base, extract business constraints for each verification dimension from the dynamic rule knowledge base, and identify rule version status, which includes existing rules and newly added rules;

[0091] If it is an existing rule, the document data object is pattern matched with the business constraints of the corresponding dimension. If the business constraints are met, the rule is considered compliant. If it is a new rule, the compatibility of the logical expression of the business constraints and the input parameters is checked. If a compatibility conflict exists, the conflicting business constraints are frozen.

[0092] The multi-dimensional compliance verification stage also includes:

[0093] When a business constraint is detected as a newly added rule, check whether the business constraint covers all data dimensions. If any dimension is missing, suspend the business constraint from taking effect.

[0094] In the above embodiment, multi-dimensional compliance verification is performed on the document data object to ensure that the data complies with business rules and prevent non-compliant data from flowing into subsequent links. A dynamic rule knowledge base is built to support automatic updating and version management of rules to ensure the adaptability and continued effectiveness of business rules. Pattern matching is performed on each dimension of the document data object to automatically determine whether it complies with business constraints. For newly added rules, the system will automatically check the compatibility of its logical expression with the input parameters to ensure that the rules will not conflict or misjudge.

[0095] The smart verification process stage also includes:

[0096] Based on a review rule set trained on big data, pattern matching is performed on compliant data objects to extract abnormal features. Combined with a multi-level verification mechanism, rule constraint verification is performed on compliant data objects, and a credibility score is calculated for compliant data objects. The pass / fail criteria are then set based on preset score thresholds.

[0097] If the verification fails, the process will be blocked and an abnormal pattern analysis report will be generated. If the verification passes, the compliant data object will be output to the next stage.

[0098] In the above embodiment, a multi-level verification mechanism is used to identify abnormal patterns and generate credibility scores to ensure the reliability of data quality. The machine learning algorithm is used to perform pattern matching on the data, which can accurately identify abnormal features in the data. The credibility score of the data is calculated through a review rule set trained with big data, and it is determined whether the data passes the verification based on the preset threshold. If the data fails the verification, the system will automatically generate an abnormal pattern analysis report to assist operators in identifying the root cause of the problem and improving data governance capabilities.

[0099] The audit task scheduling phase also includes:

[0100] Quantify the business complexity, risk factor, and timeliness priority of compliance data objects, generate task scheduling weight parameters, and match compliance data objects to the optimal review path based on the task scheduling weight parameters;

[0101] Generate audit tasks based on compliance data objects, assign audit tasks to adapted audit terminals, and after the audit is completed, output the final processing results and archive them in the business database.

[0102] In the above embodiment, intelligent task routing is performed through verified compliant data objects to ensure the reasonable allocation of audit tasks and improve audit efficiency. The system will calculate the task scheduling weight parameters based on business complexity, risk factor and time priority to ensure the best allocation plan for audit tasks. It uses intelligent scheduling algorithms to reasonably allocate audit tasks, avoid overloading of certain audit terminal tasks, and improve overall processing capabilities. If the audit task exceeds the preset time threshold, the system will automatically trigger the task reallocation strategy to ensure a smooth audit process and improve business processing efficiency.

[0103] Specifically, the quantitative calculation of the business complexity, risk factor, and timeliness priority of the compliant data objects to generate task scheduling weight parameters includes:

[0104] Extracting the data types contained in the compliant data object and the key data fields contained in each data type;

[0105] Retrieve the preset correlation value between the key data field contained in each data type and other key data fields of the same data type within the data type;

[0106] Obtaining a first complexity coefficient corresponding to each data type using a preset correlation value between a key data field contained in each data type and other key data fields of the same data type within the data type to which the key data field belongs;

[0107] The first complexity coefficient is obtained by the following formula:

[0108]

[0109] Among them, R 01 Indicates the first complex coefficient corresponding to each data type; n indicates the number of key data fields; L bi Indicates the standard deviation of the correlation between the i-th key data field and other key data fields of the same data type within the data type in which it is located; L maxi Indicates the maximum value of the correlation between the i-th key data field and other key data fields of the same data type within the data type in which it is located; L mini Indicates the minimum value of the correlation between the i-th key data field and other key data fields of the same data type within the data type in which it is located; x i Indicates the number of maximum correlation values ​​corresponding to the i-th key data field; y iIndicates the number of minimum correlation values ​​corresponding to the i-th key data field;

[0110] Retrieve the preset correlation values ​​between the key data fields contained in each data type and other data types;

[0111] Obtaining a second complexity coefficient corresponding to each data type by utilizing the first complexity coefficient corresponding to each data type in combination with a preset correlation value between the key data field contained in each data type and other data types;

[0112] The second complexity coefficient is obtained by the following formula:

[0113]

[0114] Among them, R 02 Indicates the second complex coefficient corresponding to each data type; R 01 Indicates the first complex coefficient corresponding to each data type; n indicates the number of key data fields; L bgi Indicates the standard deviation of the preset correlation between the i-th key data field and other data types; L gmaxi Indicates the maximum value of the correlation between the i-th key data field and other data types; L gmini Indicates the minimum value of the correlation between the i-th key data field and other data types; A i Indicates the number of maximum values ​​of the correlation between the i-th key data field and other data types; B i Indicates the number of minimum values ​​of the correlation between the i-th key data field and other data types; Z 01i Indicates the number of other data types included when the correlation value between the i-th key data field and other data types is the maximum; Z 02i Indicates the number of other data types included when the correlation value between the i-th key data field and other data types is the maximum;

[0115] The task scheduling weight parameter is generated by utilizing the risk coefficient corresponding to the compliant data object and the quantitative value of the timeliness priority in combination with the second complex coefficient.

[0116] The technical effect of the above technical solution is that by quantitatively calculating the business complexity, risk factor, and timeliness priority of compliant data objects and generating task scheduling weight parameters, this solution enables more refined task scheduling. Resources and processing order can be intelligently allocated based on the complexity, potential risk, and urgency of the data, thereby improving overall processing efficiency and response speed. By calculating the first and second complexity coefficients, the above technical solution comprehensively considers the correlation within and between different data types, helping the system more accurately assess data complexity and processing difficulty. Based on this, the system can prioritize data objects with lower complexity and lower risk but higher timeliness requirements, thereby improving overall data processing efficiency. By performing quantitative analysis on compliant data objects, this solution not only improves data processing efficiency but also enhances data compliance management. Based on the quantitative results, the system can allocate more attention and processing resources to high-risk or high-complexity data objects, ensuring data processing compliance and accuracy. The quantitative calculation method and formula design in the solution are flexible and scalable. As business evolves and changes, the system can adjust the quantitative standards for correlation, risk factor, and timeliness priority based on actual conditions to adapt to new data processing needs. This solution provides intelligent decision-making support for data processing. Through quantitative analysis and the generation of weighted parameters, the system can provide managers with intuitive and accurate data processing priorities and resource allocation recommendations, helping them make more scientific and reasonable decisions.

[0117] At the same time, through multi-dimensional correlation values ​​and complex calculation formulas, we can comprehensively and accurately assess the complexity within and between data types, avoiding the one-sidedness of single-dimensional assessments. The generated task scheduling weight parameters take business complexity into account, allowing for the rational allocation of resources based on the complexity of data objects during task scheduling, improving the efficiency and accuracy of task processing. For compliant data objects containing multiple data types and complex inter-data relationships, this solution can effectively quantify complexity, making it suitable for complex data processing scenarios.

[0118] On the other hand, the number of data types and key data fields in compliant data objects may continue to increase. This solution's complexity calculation method is based on an independent analysis of each key data field and its associated relationships. When a new data type or field is added, the solution simply extracts and calculates the relevant correlation values ​​according to established rules and then substitutes them into the formula for calculation. This eliminates the need for major changes to the overall architecture and is well-suited to growing data size and complexity. When new data types or relationships emerge, the solution calculates complexity using preset correlation values. This allows for the flexibility of adding preset values ​​for new correlation patterns to the system, allowing new relationships to be incorporated into the complexity calculation, ensuring the solution's scalability to adapt to changing business needs. During the complexity coefficient calculation process, relevant calculations for each key data field (such as the statistical characteristics of the correlation values ​​and the number of occurrences) can be performed in parallel. This allows the use of multi-core processors or distributed computing environments when processing large-scale data objects, allowing for simultaneous calculations of multiple key data fields. This significantly reduces overall calculation time and improves the efficiency of business complexity assessment. The solution retrieves and uses correlation values ​​based on preset configurations. During actual calculations, only the corresponding values ​​need to be retrieved according to the rules. This avoids repeated calculations and unnecessary data processing, reduces the waste of computing resources, and further improves processing efficiency. Because complexity assessment can generate accurate task scheduling weight parameters, the system can rationally allocate computing resources based on the complexity of different tasks. For tasks with lower complexity, fewer computing resources are allocated to complete the processing; for complex tasks, more resources are allocated to ensure efficient completion. This avoids over- or under-allocation of resources and improves overall computing resource utilization.

[0119] In summary, this technical solution achieves refined task scheduling and efficient processing of compliant data objects through quantitative calculation and generation of weight parameters, improves data processing efficiency and compliance management level, and provides strong support for enterprise data management and business development.

[0120] Specifically, the risk coefficient corresponding to the compliant data object and the quantitative value of the timeliness priority are combined with the second complex coefficient to generate the task scheduling weight parameter, including:

[0121] Identifying risk events contained in the compliance data objects;

[0122] Performing a risk probability analysis on the risk events to obtain the corresponding probability of occurrence of each risk event;

[0123] The risk factor corresponding to each compliant data object is obtained by performing a weighted average using the probability of occurrence of the risk event and the weight parameter corresponding to the risk event;

[0124] Identify key time nodes in the processing of compliant data objects and obtain the time difference between each key time node and the current time;

[0125] Identify the incidence of risk events corresponding to the key time nodes;

[0126] The quantitative value of the timeliness priority is obtained by using the occurrence rate of the risk events corresponding to the key time nodes and the time difference between each key time node and the current time;

[0127] The quantified value of the timeliness priority is obtained by the following formula:

[0128]

[0129] Among them, F represents the quantitative value of time priority; m represents the number of key time nodes; P i represents the occurrence rate of risk events corresponding to the i-th key time node; T i Indicates the time difference between the i-th key time node and the current time; T c Indicates the preset time difference reference value;

[0130] Retrieve the second complex coefficient;

[0131] The risk coefficient and the quantified value of the timeliness priority corresponding to the compliant data object are combined with the second complex coefficient to generate an initial task scheduling weight parameter; wherein the initial task scheduling weight parameter is obtained by the following formula:

[0132]

[0133] Among them, W ctask represents the initial task scheduling weight parameter; k represents the number of data types contained in the compliant data object; R 02i represents the second complexity coefficient corresponding to the i-th data type; F represents the quantitative value of the timeliness priority; U represents the risk coefficient corresponding to the compliance data object;

[0134] The initial task scheduling weight parameter corresponding to each compliant data object is standardized to generate the task scheduling weight parameter corresponding to each compliant data object.

[0135] The technical effects of the above technical solution are as follows: First, risk events in compliant data objects are identified, the probability of each risk event is analyzed, and then a weighted average is taken using the risk event probability and its corresponding weight parameter to determine the risk coefficient corresponding to each compliant data object, thereby measuring the degree of risk faced by the data object. Key time nodes in the data object processing process are identified, the time difference between each key time node and the current time is calculated, and the risk event occurrence rate corresponding to the key time node is determined. A quantitative value of the timeliness priority is calculated using a specific formula (comprehensively considering the number of key time nodes, the risk event occurrence rate at each node, the time difference, and a preset time difference reference value), reflecting the time and risk urgency of data processing. The previously calculated second complexity coefficient is retrieved, combined with the risk coefficient and the quantitative value of the timeliness priority, and an initial task scheduling weight parameter is generated according to a corresponding formula (taking into account factors such as the number of data types and the second complexity coefficient of each data type), thereby preliminarily determining the weight of data object task scheduling. The initial task scheduling weight parameter is standardized to eliminate the dimensionality effects of different data types, resulting in a final task scheduling weight parameter corresponding to each compliant data object, which is used to guide task scheduling.

[0136] By comprehensively considering the risk factor, timeliness priority, and secondary complexity factor of each compliance data object, this solution generates a comprehensive and accurate initial task scheduling weight parameter. This weight parameter not only reflects the urgency and importance of data processing, but also takes into account the complexity and processing difficulty of the data itself, providing a scientific basis for task scheduling. Using this weight parameter, the system automatically prioritizes compliance data objects, ensuring that critical and urgent tasks are prioritized. The solution's identification and analysis of risk events, as well as the calculation of risk factors, help enterprises better identify and manage potential risks. By quantifying the probability of risk occurrence and the corresponding weight, enterprises can more intuitively understand the risk level of each compliance data object and implement appropriate risk management measures to mitigate the likelihood and impact of risk. By identifying key time nodes and calculating the time difference, and combining the occurrence rate of risk events to derive a quantitative value for timeliness priority, the solution fully considers the timeliness requirements of data processing. This helps enterprises ensure that critical data is processed within the specified timeframe, avoiding issues caused by delays. Furthermore, by optimizing task scheduling, enterprises can improve data processing efficiency and reduce unnecessary waiting and delays. By standardizing the initial task scheduling weight parameters, this solution ensures comparability and consistency across all compliant data objects. This facilitates data sharing and collaboration across different departments or teams within an enterprise, improving overall work efficiency and synergy.

[0137] At the same time, multi-dimensional comprehensive calculations acquire data from multiple perspectives, such as risk events and time nodes, and separately calculate the risk coefficient, quantitative value of timeliness priority, and second complexity coefficient. These are then integrated so that the weight parameters fully reflect the business characteristics and processing requirements of the data object, reducing errors caused by one-sided assessments and improving the accuracy of the weight parameters in guiding actual task scheduling. Operations such as risk event identification and key time node determination have a certain degree of independence. When business changes introduce new risk event types or key time nodes, only relevant data and calculations need to be supplemented according to established rules, without the need for major adjustments to the overall architecture. This allows the system to adapt to changes in data and processing requirements brought about by business development, ensuring the scalability of the weight parameter acquisition solution.

[0138] In summary, this technical solution achieves scientific assessment and optimization of task scheduling by comprehensively considering the risk factor, timeliness priority, and secondary complexity factor of compliant data objects. This helps enterprises improve the efficiency and accuracy of data processing, reduce risk levels, and enhance overall business operations and competitiveness.

[0139] See also Figure 3 The document information entry system based on big data processing applies the above-mentioned document information entry method based on big data processing, including:

[0140] Heterogeneous data source parsing module, used for:

[0141] Extract original document data stored in various media carriers, parse document data in different formats, and generate standardized data input streams;

[0142] Intelligent data mapping module for:

[0143] Receive the parsed original document data, align the data according to the template framework, and generate a structured document data object;

[0144] Risk signature screening module for:

[0145] Scan document data objects for risk characteristics, analyze their characteristics, extract key information elements, and identify potential sensitive elements based on these key information elements;

[0146] Multi-dimensional compliance verification module, used for:

[0147] Extract business constraints for each verification dimension from the dynamic rule knowledge base, identify the rule version status, and perform multi-dimensional compliance verification on document data objects based on the business constraints;

[0148] Intelligent verification process module for:

[0149] Based on the review rule set trained on big data, multi-level verification is carried out on compliant data objects, abnormal features are extracted, rule constraints are verified on compliant data objects, and the credibility score of compliant data objects is calculated;

[0150] Audit task scheduling module, used to:

[0151] Quantify the business complexity, risk factor, and timeliness priority of compliance data objects, generate task scheduling weight parameters, and match compliance data objects to the optimal review path based on the task scheduling weight parameters;

[0152] Assign audit tasks to the audit terminal, and archive compliant data objects to the business database after the audit is completed.

[0153] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A document information entry method based on big data processing, characterized in that: include: Collecting original document data and identifying compliant data objects in the original document data, generating task scheduling weight parameters for the compliant data objects, wherein the process of determining the task scheduling weight parameters includes: retrieving a correlation value based on a data type and a key data field of the compliant data object, obtaining a first complexity coefficient and a second complexity coefficient based on the correlation value, and generating the task scheduling weight parameter in combination with a risk coefficient and a timeliness priority quantified value; Match each compliant data object to the optimal audit path based on the task scheduling weight parameters. After the compliant data object is audited through the optimal audit path, the final processing results are output and archived to the business database. Wherein, in the process of determining the task scheduling weight parameter, the data types contained in the compliance data object and the key data fields contained in each data type are extracted; Retrieve the preset correlation value between the key data field contained in each data type and other key data fields of the same data type within the data type; Obtaining a first complexity coefficient corresponding to each data type using a preset correlation value between a key data field contained in each data type and other key data fields of the same data type within the data type to which the key data field belongs; Retrieve the preset correlation values ​​between the key data fields contained in each data type and other data types; Obtaining a second complexity coefficient corresponding to each data type by utilizing the first complexity coefficient corresponding to each data type in combination with a preset correlation value between the key data field contained in each data type and other data types; Identifying risk events contained in the compliance data objects; Performing a risk probability analysis on the risk events to obtain the corresponding probability of occurrence of each risk event; The risk factor corresponding to each compliant data object is obtained by performing a weighted average using the probability of occurrence of the risk event and the weight parameter corresponding to the risk event; Identify key time nodes in the processing of compliant data objects and obtain the time difference between each key time node and the current time; Identify the incidence of risk events corresponding to the key time nodes; The quantitative value of the timeliness priority is obtained by using the occurrence rate of the risk events corresponding to the key time nodes and the time difference between each key time node and the current time; The task scheduling weight parameter is generated by utilizing the risk coefficient corresponding to the compliance data object and the quantitative value of the timeliness priority in combination with the second complex coefficient.

2. The document information entry method based on big data processing according to claim 1, characterized in that: include: Extract original document data stored in diverse media carriers, including database clusters, optical character recognition files, electronic communication attachments, and IoT device record streams, align the original document data to a template framework, and generate structured document data objects.

3. The document information entry method based on big data processing according to claim 1, characterized in that: include: Based on big data mining and combined with knowledge graphs, document data objects are scanned for risk characteristics to identify sensitive elements in the preset risk characteristic library; Analyze the characteristics of document data objects based on big data mining technology, extract key information elements from the document, compare the extracted key information elements with the preset risk feature library, and identify possible sensitive elements; If sensitive elements are detected, the process will be blocked, and the risk category and impact scope will be determined based on association rule analysis. Real-time warnings will be issued and multi-channel risk notifications will be pushed to the monitoring terminal.

4. The document information entry method based on big data processing according to claim 1, characterized in that: include: Perform multi-dimensional compliance verification on document data objects. If any dimension verification fails, the process is blocked and a visual diagnostic report is generated to assist operators in locating the root cause of the anomaly. If the verification succeeds, the compliant data object is output. Build a dynamic rule knowledge base, extract business constraints for each verification dimension from the dynamic rule knowledge base, and identify rule version status, which includes existing rules and newly added rules; If it is an existing rule, the document data object is pattern matched with the business constraints of the corresponding dimension. If the business constraints are met, the rule is considered compliant. If it is a new rule, the compatibility of the logical expression of the business constraint and the input parameters is checked. If there is a compatibility conflict, the conflicting business constraint is frozen. When a business constraint is detected as a newly added rule, check whether the business constraint covers all data dimensions. If any dimension is missing, suspend the business constraint from taking effect.

5. The document information entry method based on big data processing according to claim 1, characterized in that: include: Based on a set of review rules trained on big data, multi-level verification is performed on compliant data objects to identify abnormal patterns and generate credibility scores; Based on a review rule set trained on big data, pattern matching is performed on compliant data objects to extract abnormal features. Combined with a multi-level verification mechanism, rule constraint verification is performed on compliant data objects, and a credibility score is calculated for compliant data objects. The pass / fail criteria are then set based on preset score thresholds. If the verification fails, the process will be blocked and an abnormal pattern analysis report will be generated. If the verification passes, the compliant data object will be output to the next stage.

6. The document information entry method based on big data processing according to claim 1, characterized in that: include: Intelligent task routing is performed on compliant data objects that have passed verification, matching each compliant data object to the optimal review path based on multiple criteria such as business complexity, risk factor, and timeliness priority. Quantify the business complexity, risk factor, and timeliness priority of compliance data objects, generate task scheduling weight parameters, and match compliance data objects to the optimal review path based on the task scheduling weight parameters; Generate audit tasks based on compliance data objects, assign audit tasks to adapted audit terminals, and after the audit is completed, output the final processing results and archive them in the business database.

7. The document information entry method based on big data processing according to claim 1, characterized in that: The quantitative value of the timeliness priority is obtained by the following formula: Among them, F represents the quantitative value of time priority; m represents the number of key time nodes; P i represents the occurrence rate of risk events corresponding to the i-th key time node; T i Indicates the time difference between the i-th key time node and the current time; T c Indicates the preset time difference reference value.

8. The document information entry method based on big data processing according to claim 7, characterized in that: Generating a task scheduling weight parameter by using the risk coefficient corresponding to the compliance data object and the quantitative value of the timeliness priority in combination with the second complex coefficient also includes: Retrieve the second complex coefficient; The risk coefficient corresponding to the compliance data object and the quantitative value of the timeliness priority are combined with the second complex coefficient to generate an initial task scheduling weight parameter; wherein the initial task scheduling weight parameter is obtained by the following formula: Among them, W ctask represents the initial task scheduling weight parameter; k represents the number of data types contained in the compliant data object; R 02i represents the second complexity coefficient corresponding to the i-th data type; F represents the quantitative value of the timeliness priority; U represents the risk coefficient corresponding to the compliance data object; The initial task scheduling weight parameter corresponding to each compliant data object is standardized to generate the task scheduling weight parameter corresponding to each compliant data object.

9. A document information entry system based on big data processing, applied to a document information entry method based on big data processing according to any one of claims 1 to 8, characterized in that: include: Heterogeneous data source parsing module, used for: Extract original document data stored in various media carriers, parse document data in different formats, and generate standardized data input streams; Intelligent data mapping module for: Receive the parsed original document data, align the data according to the template framework, and generate a structured document data object; Risk signature screening module for: Scan document data objects for risk characteristics, analyze their characteristics, extract key information elements, and identify potential sensitive elements based on these key information elements; Multi-dimensional compliance verification module, used for: Extract business constraints for each verification dimension from the dynamic rule knowledge base, identify the rule version status, and perform multi-dimensional compliance verification on document data objects based on the business constraints; Intelligent verification process module for: Based on the review rule set trained on big data, multi-level verification is carried out on compliant data objects, abnormal features are extracted, rule constraints are verified on compliant data objects, and the credibility score of compliant data objects is calculated; Audit task scheduling module, used to: Quantify the business complexity, risk factor and time priority of compliance data objects, generate task scheduling weight parameters, match compliance data objects to the optimal audit path based on the task scheduling weight parameters, assign audit tasks to audit terminals, and archive compliance data objects to the business database after the audit is completed.

Citation Information

Patent Citations

  • Receipt information input method and device based on data processing, and storage medium

    CN109739957A

  • Method applied to enterprise OA collaborative workflow optimization

    CN118761745A