Big data processing-based document information input method and system

By collecting and identifying compliant data objects in the document information entry system, combining big data mining and knowledge graph for risk analysis, and dynamically updating business rules, the problem of difficulty in in-depth risk analysis and adapting to changes in business rules in the existing technology is solved, and efficient document data processing and compliance inspection are achieved.

CN119990719AActive Publication Date: 2025-05-13SHENZHEN TEWEI KECHUANG INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510477952.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

In the process of document information entry, it is difficult for the existing technology to make full use of big data mining technology for in-depth risk analysis, resulting in false positives or missed reports, and it is difficult to dynamically adapt to changes in business rules, resulting in lagging in rule updates or incomplete verification.

Method used

By collecting and identifying compliant data objects in document data, generating task scheduling weight parameters, combining big data mining and knowledge graphs to scan risk characteristics, dynamically update business rules, and optimize audit paths.

Benefits of technology

It realizes efficient document data processing and risk analysis, improves the accuracy and timeliness of compliance inspections, and optimizes audit efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990719A_ABST
    Figure CN119990719A_ABST
Patent Text Reader

Abstract

The invention discloses a document information input method and system based on big data processing, and belongs to the technical field of document information processing. According to the method, deep risk feature screening is performed on document data by combining big data mining and a knowledge graph, potential sensitive elements can be accurately identified, risk categories and influence ranges are analyzed based on association rules, automatic early warning and real-time risk notification are realized, a dynamic rule knowledge base is adopted, multi-dimensional compliance verification is supported, and the risk assessment efficiency is improved. Service constraint conditions can be updated in real time, the flexibility and adaptability of rules are ensured, the accuracy and timeliness of compliance check are improved, pattern matching and multi-level check are performed on compliance data objects through an intelligent check process and a recheck rule set trained by big data, the preciseness and reliability of data check are ensured, and the data check efficiency is improved. Through comprehensive quantitative calculation of the business complexity, the risk coefficient and the timeliness priority, the task scheduling weight parameter is generated, the receipt auditing path is optimized, and the auditing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document information processing, and in particular to a document information entry method and system based on big data processing. Background Art

[0002] In the context of modern data management and business process automation, various enterprises, financial institutions and government departments need to process a large amount of document information, such as invoices, contracts, reimbursement documents, bank statements, etc. These documents usually come from a wide range of sources and in various formats, including both structured data and unstructured or semi-structured data. With the increasing compliance and regulatory requirements of enterprises, document data needs to undergo strict risk screening and compliance verification during the input process. At present, existing technologies still have some limitations, such as: Traditional risk control methods often rely on preset rules and cannot fully utilize big data mining technology for in-depth risk analysis, which is prone to false positives or omissions. Some existing systems use fixed rule sets for compliance checks, which are difficult to dynamically adapt to changing business rules, resulting in delayed rule updates or incomplete verification. The allocation of audit tasks usually relies on manual judgment or simple fixed rules, and fails to perform intelligent scheduling based on business complexity, risk level and audit priority, affecting audit efficiency and accuracy. Summary of the invention

[0003] The purpose of the present invention is to provide a document information entry method and system based on big data processing to solve the problems raised in the above background technology.

[0004] To achieve the above object, the present invention provides the following technical solution: a document information entry method based on big data processing, comprising: Collecting original document data and identifying compliant data objects in the original document data, generating task scheduling weight parameters for the compliant data objects, wherein the process of determining the task scheduling weight parameters includes: retrieving a correlation value based on a data type and a key data field of the compliant data object, obtaining a first complexity coefficient and a second complexity coefficient based on the correlation value, and generating the task scheduling weight parameters in combination with a risk coefficient and a timeliness priority quantified value; Based on the task scheduling weight parameters, each compliant data object is matched to the optimal audit path. After the compliant data object is audited through the optimal audit path, the final processing result is output and archived to the business database.

[0005] Furthermore, it includes: extracting original document data stored in diverse media carriers, aligning the original document data to a template framework and generating a structured document data object, wherein the diverse media carriers include database clusters, optical character recognition files, electronic communication attachment documents, and IoT device record streams.

[0006] Further, it includes: scanning the risk characteristics of document data objects based on big data mining and combining with knowledge graphs to identify sensitive elements in the preset risk characteristic library; Based on big data mining technology, feature analysis is performed on document data objects to extract key information elements from the document. The extracted key information elements are compared with the preset risk feature library to identify possible sensitive elements. If sensitive elements are detected, the process will be blocked, and the risk category and impact scope will be determined based on association rule analysis. Real-time warnings will be issued and multi-channel risk notifications will be pushed to the monitoring terminal.

[0007] Furthermore, it includes: performing multi-dimensional compliance verification on the document data object. If any dimension verification fails, the process is blocked and a visual diagnosis report is generated to assist the operator in locating the root cause of the abnormality. If the verification succeeds, the compliant data object is output; Build a dynamic rule knowledge base, extract business constraints of each verification dimension from the dynamic rule knowledge base, and identify the rule version status, which includes existing rules and newly added rules; If it is an existing rule, the document data object is pattern matched with the business constraints of the corresponding dimension; if the business constraints are met, it is determined to be compliant; if it is a new rule, the compatibility of the logical expression of the business constraints and the input parameters is checked, and if there is a compatibility conflict, the conflicting business constraints are frozen; When it is detected that the business constraint is a newly added rule, check whether the business constraint covers all data dimensions. If there are missing dimensions, suspend the business constraint from taking effect.

[0008] Further, it includes: implementing multi-level verification of compliant data objects based on a review rule set trained on big data, identifying abnormal patterns and generating credibility scores; Based on the review rule set trained with big data, pattern matching is performed on compliant data objects to extract abnormal features. In combination with a multi-level verification mechanism, rule constraint verification is performed on compliant data objects, the credibility score of compliant data objects is calculated, and the judgment criteria for passing or failing the verification are set based on the preset score threshold; If the verification fails, the process will be blocked and an abnormal pattern analysis report will be generated. If the verification passes, the compliant data object will be output to the next stage.

[0009] Furthermore, it includes: intelligent task routing for compliant data objects that have passed verification, matching each compliant data object to the optimal audit path based on business complexity, risk factor and timeliness priority; Quantify the business complexity, risk factor and timeliness priority of compliant data objects, generate task scheduling weight parameters, and match compliant data objects to the optimal audit path based on the task scheduling weight parameters; Generate audit tasks based on compliance data objects, assign audit tasks to compatible audit terminals, and after the audit is completed, output the final processing results and archive them in the business database.

[0010] Furthermore, the quantitative calculation of the business complexity, risk factor and timeliness priority of the compliant data object to generate the task scheduling weight parameter includes: Extracting the data types contained in the compliant data object and the key data fields contained in each data type; Retrieving a preset correlation value between the key data field contained in each data type and other key data fields of the same data type in the data type; Obtaining a first complexity coefficient corresponding to each data type by using a preset correlation value between a key data field contained in each data type and other key data fields of the same data type within the data type in which the key data field is located; The first complexity coefficient is obtained by the following formula: Among them, R 01 Indicates the first complexity coefficient corresponding to each data type; n indicates the number of key data fields; L bi Indicates the standard deviation of the correlation between the i-th key data field and other key data fields of the same data type in the data type; L maxi Indicates the maximum value of the correlation between the i-th key data field and other key data fields of the same data type in the data type in which it is located; L mini Indicates the minimum value of the correlation between the i-th key data field and other key data fields of the same data type in the data type; x i Indicates the number of maximum correlation values ​​corresponding to the i-th key data field; y i Indicates the number of minimum values ​​of the correlation degree corresponding to the i-th key data field; Retrieve the preset correlation value between the key data fields contained in each data type and other data types; Obtaining a second complexity coefficient corresponding to each data type by using the first complexity coefficient corresponding to each data type in combination with a preset correlation value between a key data field contained in each data type and other data types; The second complexity coefficient is obtained by the following formula: Among them, R 02 Indicates the second complex coefficient corresponding to each data type; R 01 Indicates the first complexity coefficient corresponding to each data type; n indicates the number of key data fields; L bgi Indicates the standard deviation of the preset correlation between the i-th key data field and other data types; L gmaxi Indicates the maximum value of the correlation between the i-th key data field and other data types; L gmini Indicates the minimum value of the correlation between the i-th key data field and other data types; A i Indicates the number of maximum values ​​of the correlation between the i-th key data field and other data types; B i Indicates the number of minimum values ​​of the correlation between the i-th key data field and other data types; Z 01i Indicates the number of other data types included in the maximum correlation value when the i-th key data field has the maximum correlation value with other data types; Z 02i Indicates the number of other data types included in the minimum correlation value when the correlation value between the i-th key data field and other data types is the maximum; The risk coefficient corresponding to the compliant data object and the quantified value of the timeliness priority are combined with the second complex coefficient to generate a task scheduling weight parameter.

[0011] Furthermore, the risk coefficient corresponding to the compliant data object and the quantitative value of the timeliness priority are combined with the second complex coefficient to generate a task scheduling weight parameter, including: Identifying risk events contained in the compliance data object; Performing risk probability analysis on the risk events to obtain the corresponding probability of occurrence of each risk event; The risk factor corresponding to each compliant data object is obtained by performing weighted averaging using the probability of occurrence of the risk event and the weight parameter corresponding to the risk event; Identify key time nodes in the processing of compliant data objects and obtain the time difference between each key time node and the current time; Identify the occurrence rate of risk events corresponding to the key time nodes; The quantified value of the timeliness priority is obtained by using the occurrence rate of the risk events corresponding to the key time nodes and the time difference between each key time node and the current time; The quantified value of the timeliness priority is obtained by the following formula: Among them, F represents the quantitative value of timeliness priority; m represents the number of key time nodes; P i represents the occurrence rate of risk events corresponding to the i-th key time node; T i represents the time difference between the i-th key time node and the current time; T c Indicates the preset time difference reference value; Retrieve the second complex coefficient; The risk coefficient corresponding to the compliant data object and the quantitative value of the timeliness priority are combined with the second complex coefficient to generate an initial task scheduling weight parameter; wherein the initial task scheduling weight parameter is obtained by the following formula: Among them, W ctask represents the initial task scheduling weight parameter; k represents the number of data types contained in the compliant data object; R 02i represents the second complexity coefficient corresponding to the i-th data type; F represents the quantitative value of the timeliness priority; U represents the risk coefficient corresponding to the compliance data object; The initial task scheduling weight parameter corresponding to each compliant data object is standardized to generate the task scheduling weight parameter corresponding to each compliant data object.

[0012] Furthermore, the document information entry system based on big data processing is applied in the above-mentioned document information entry method based on big data processing, including: Heterogeneous data source parsing module, used for: Extract original document data stored in various media carriers, parse document data in different formats, and generate standardized data input streams; Intelligent data mapping module for: Receive the parsed original document data, align the data according to the template framework, and generate a structured document data object; Risk Signature Screening Module for: Scan the document data objects for risk characteristics, analyze the characteristics of the document data objects, extract key information elements, and identify possible sensitive elements based on the key information elements; Multi-dimensional compliance verification module, used for: Extract business constraints of each verification dimension from the dynamic rule knowledge base and identify the rule version status, and perform multi-dimensional compliance verification on document data objects based on the business constraints.

[0013] Furthermore, the document information entry system based on big data processing also includes: Intelligent verification process module for: Based on the review rule set trained on big data, multi-level verification is carried out on compliant data objects, abnormal features are extracted, rule constraint verification is performed on compliant data objects, and credibility scores of compliant data objects are calculated; Audit task scheduling module, used to: Quantify the business complexity, risk factor and timeliness priority of compliant data objects, generate task scheduling weight parameters, and match compliant data objects to the optimal audit path based on the task scheduling weight parameters; Assign audit tasks to audit terminals, and archive compliant data objects to the business database after the audit is completed.

[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention adopts heterogeneous data source parsing technology, which can efficiently extract original document data stored in databases, OCR files, electronic communication accessories and IoT device record streams, realize automatic parsing and formatting of multi-source data, and convert the parsed unstructured or semi-structured document data into standardized structured data objects through intelligent data mapping, thereby improving the consistency and accuracy of data processing and reducing manual intervention.

[0015] 2. The present invention combines big data mining and knowledge graphs to conduct in-depth risk feature screening of document data, can accurately identify potential sensitive factors, and analyze risk categories and impact scope based on association rules, to achieve automatic early warning and real-time risk notification, adopt a dynamic rule knowledge base, support multi-dimensional compliance verification, and can update business constraints in real time to ensure the flexibility and adaptability of rules, improve the accuracy and timeliness of compliance checks, and through an intelligent verification process, use the review rule set trained by big data to perform pattern matching and multi-level verification on compliant data objects, and generate credibility scores to ensure the rigor and reliability of data audits.

[0016] 3. The present invention generates task scheduling weight parameters through comprehensive quantitative calculation of business complexity, risk factor and timeliness priority, optimizes the document review path, improves review efficiency, adopts a dynamic load balancing algorithm, realizes intelligent task allocation, ensures optimal utilization of review resources, reduces manual review pressure, improves overall business processing capabilities, supports automatic archiving and recording of review tasks, ensures the integrity and traceability of data tracking, and meets regulatory and audit requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the document information input method of the present invention; Figure 2 This is a schematic diagram of the document information entry process of the present invention; Figure 3 It is a schematic diagram of the document information entry system module of the present invention. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] See also Figure 1-Figure 2 , the present invention provides the following technical solutions: The document information entry method based on big data processing includes: Collecting original document data and identifying compliant data objects in the original document data, generating task scheduling weight parameters for the compliant data objects, wherein the process of determining the task scheduling weight parameters includes: retrieving a correlation value based on a data type and a key data field of the compliant data object, obtaining a first complexity coefficient and a second complexity coefficient based on the correlation value, and generating the task scheduling weight parameters in combination with a risk coefficient and a timeliness priority quantified value; Based on the task scheduling weight parameters, each compliant data object is matched to the optimal audit path. After the compliant data object is audited through the optimal audit path, the final processing result is output and archived to the business database.

[0020] The specific steps include: Heterogeneous data source analysis phase: extracting original document data stored in a variety of media carriers, including database clusters, optical character recognition files, electronic communication attachment documents, and IoT device record streams; Intelligent data mapping stage: align the parsed original document data to the template framework and generate structured document data objects; Risk feature screening stage: Based on big data mining and combined with knowledge graphs, risk feature scanning is performed on document data objects to identify sensitive elements in the preset risk feature library; Multi-dimensional compliance verification stage: Perform multi-dimensional compliance verification on the document data object. If any dimension verification fails, the process will be blocked and a visual diagnosis report will be generated to assist operators in locating the root cause of the abnormality. If the verification is successful, the compliant data object will be output; Intelligent verification process stage: Based on the review rule set trained based on big data, multi-level verification is carried out on compliant data objects to identify abnormal patterns and generate credibility scores; Audit task scheduling stage: Intelligent task routing is performed on compliant data objects that have passed verification, and each compliant data object is matched to the optimal audit path based on business complexity, risk factor, and time priority.

[0021] In the above embodiment, through intelligent data mapping, the parsed data is automatically aligned to the preset template framework, reducing the workload of manual adjustment and the input error rate. By using big data mining and knowledge graphs, it is possible to efficiently detect risk features in document data, accurately identify sensitive information, and improve risk management capabilities. Through a dynamic rule knowledge base, the system can adapt to changes in business rules, achieve accurate rule constraint checks, and improve the accuracy of compliance judgments. By using a review rule set based on big data training, it can automatically identify abnormal patterns, calculate credibility scores, and ensure data quality and consistency. Audit task scheduling uses an intelligent routing algorithm that combines business complexity, risk factors, and timeliness priorities to ensure efficient and reasonable allocation of audit tasks and optimize business processes.

[0022] The risk profile screening phase also includes: Based on big data mining technology, feature analysis is performed on document data objects to extract key information elements from the document. The extracted key information elements are compared with the preset risk feature library to identify possible sensitive elements. If sensitive elements are detected, the process will be blocked, and the risk category and impact scope will be determined based on association rule analysis. Real-time warnings will be issued and multi-channel risk notifications will be pushed to the monitoring terminal.

[0023] In the above embodiment, through big data mining and knowledge graph technology, the risk characteristics of the document data object are scanned to identify potential risk factors and improve the security of document processing. The use of knowledge graph and risk characteristic library can efficiently identify abnormal information in the document, such as financial fraud, data leakage and other risks. Once sensitive factors are detected, the system will immediately trigger the process blocking and push risk notifications to the risk control terminal to ensure timely intervention. Through big data mining technology, the risk category and impact range can be analyzed to improve the accuracy of risk control.

[0024] Among them, the multi-dimensional compliance verification stage further includes: Build a dynamic rule knowledge base, extract business constraints of each verification dimension from the dynamic rule knowledge base, and identify the rule version status, which includes existing rules and newly added rules; If it is an existing rule, the document data object is matched with the business constraints of the corresponding dimension; if the business constraints are met, it is determined to be compliant; if it is a new rule, the compatibility of the logical expression of the business constraints and the input parameters is checked, and if there is a compatibility conflict, the conflicting business constraints are frozen.

[0025] The multi-dimensional compliance verification stage also includes: When it is detected that the business constraint is a newly added rule, check whether the business constraint covers all data dimensions. If there are missing dimensions, suspend the business constraint from taking effect.

[0026] In the above embodiment, multi-dimensional compliance verification is performed on the document data object to ensure that the data complies with business rules and to prevent non-compliant data from flowing into subsequent links. A dynamic rule knowledge base is built to support automatic updating and version management of rules to ensure the adaptability and continued effectiveness of business rules. Pattern matching is performed on each dimension of the document data object to automatically determine whether it complies with business constraints. For newly added rules, the system will automatically check the compatibility of its logical expression with the input parameters to ensure that the rules will not cause conflicts or misjudgments.

[0027] The smart verification process stage also includes: Based on the review rule set trained with big data, pattern matching is performed on compliant data objects to extract abnormal features. In combination with a multi-level verification mechanism, rule constraint verification is performed on compliant data objects, the credibility score of compliant data objects is calculated, and the judgment criteria for passing or failing the verification are set based on the preset score threshold; If the verification fails, the process will be blocked and an abnormal pattern analysis report will be generated. If the verification passes, the compliant data object will be output to the next stage.

[0028] In the above embodiment, a multi-level verification mechanism is used to identify abnormal patterns and generate credibility scores to ensure the reliability of data quality. The machine learning algorithm is used to perform pattern matching on the data, and abnormal features in the data can be accurately identified. The credibility score of the data is calculated through a review rule set trained with big data, and it is determined whether the data has passed the verification based on a preset threshold. If the data fails the verification, the system will automatically generate an abnormal pattern analysis report to assist operators in identifying the root cause of the problem and improving data governance capabilities.

[0029] The audit task scheduling phase also includes: Quantify the business complexity, risk factor and timeliness priority of compliant data objects, generate task scheduling weight parameters, and match compliant data objects to the optimal audit path based on the task scheduling weight parameters; Generate audit tasks based on compliance data objects, assign audit tasks to compatible audit terminals, and after the audit is completed, output the final processing results and archive them in the business database.

[0030] In the above embodiment, intelligent task routing is performed through verified compliant data objects to ensure the reasonable allocation of audit tasks and improve audit efficiency. The system will calculate the task scheduling weight parameters based on business complexity, risk factor and time priority to ensure the best allocation plan for audit tasks. It uses intelligent scheduling algorithms to reasonably allocate audit tasks, avoid overloading of certain audit terminal tasks, and improve overall processing capabilities. If the audit task exceeds the preset time threshold, the system will automatically trigger the task reallocation strategy to ensure a smooth audit process and improve business processing efficiency.

[0031] Specifically, the quantitative calculation of the business complexity, risk factor and timeliness priority of the compliant data object to generate the task scheduling weight parameter includes: Extracting the data types contained in the compliant data object and the key data fields contained in each data type; Retrieving a preset correlation value between the key data field contained in each data type and other key data fields of the same data type in the data type; Obtaining a first complexity coefficient corresponding to each data type by using a preset correlation value between a key data field contained in each data type and other key data fields of the same data type within the data type in which the key data field is located; The first complexity coefficient is obtained by the following formula: Among them, R 01 Indicates the first complexity coefficient corresponding to each data type; n indicates the number of key data fields; L bi Indicates the standard deviation of the correlation between the i-th key data field and other key data fields of the same data type in the data type; L maxi Indicates the maximum value of the correlation between the i-th key data field and other key data fields of the same data type in the data type in which it is located; L mini Indicates the minimum value of the correlation between the i-th key data field and other key data fields of the same data type in the data type; x i Indicates the number of maximum correlation values ​​corresponding to the i-th key data field; y i Indicates the number of minimum values ​​of the correlation degree corresponding to the i-th key data field; Retrieve the preset correlation value between the key data fields contained in each data type and other data types; Obtaining a second complexity coefficient corresponding to each data type by using the first complexity coefficient corresponding to each data type in combination with a preset correlation value between a key data field contained in each data type and other data types; The second complexity coefficient is obtained by the following formula: Among them, R 02 Indicates the second complex coefficient corresponding to each data type; R 01 Indicates the first complexity coefficient corresponding to each data type; n indicates the number of key data fields; L bgi Indicates the standard deviation of the preset correlation between the i-th key data field and other data types; L gmaxi Indicates the maximum value of the correlation between the i-th key data field and other data types; L gmini Indicates the minimum value of the correlation between the i-th key data field and other data types; A i Indicates the number of maximum values ​​of the correlation between the i-th key data field and other data types; B i Indicates the number of minimum values ​​of the correlation between the i-th key data field and other data types; Z 01i Indicates the number of other data types included in the maximum correlation value when the i-th key data field has the maximum correlation value with other data types; Z 02i Indicates the number of other data types included in the minimum correlation value when the correlation value between the i-th key data field and other data types is the maximum; The risk coefficient corresponding to the compliant data object and the quantified value of the timeliness priority are combined with the second complex coefficient to generate a task scheduling weight parameter.

[0032] The technical effect of the above technical solution is: by quantitatively calculating the business complexity, risk coefficient and timeliness priority of compliant data objects, generating task scheduling weight parameters, the solution can achieve more refined task scheduling. Resources and processing order can be intelligently allocated according to the complexity, potential risks and urgency of the data, thereby improving the overall processing efficiency and response speed. In the above technical solution, by calculating the first complexity coefficient and the second complexity coefficient, the correlation within the data type and between different data types is comprehensively considered, which helps the system to more accurately evaluate the complexity and processing difficulty of the data. Based on this, the system can give priority to those data objects with low complexity, low risk but high timeliness requirements, thereby improving the data processing efficiency as a whole. By quantitatively analyzing compliant data objects, the solution not only improves data processing efficiency, but also enhances data compliance management. Based on the quantitative results, the system can pay more attention and processing resources to high-risk or high-complexity data objects to ensure the compliance and accuracy of data processing. The quantitative calculation method and formula design in the solution have certain flexibility and scalability. With the development and changes of the business, the system can adjust the quantitative standards of the correlation value, risk coefficient and timeliness priority according to the actual situation to adapt to new data processing needs. The solution provides intelligent decision support for data processing. Through quantitative analysis and the generation of weight parameters, the system can provide managers with intuitive and accurate data processing priorities and resource allocation suggestions, helping them make more scientific and reasonable decisions.

[0033] At the same time, through multi-dimensional correlation values ​​and complex formula calculations, the complexity within and between data types can be comprehensively and accurately evaluated, avoiding the one-sidedness of single-dimensional evaluation. The generated task scheduling weight parameters take into account business complexity, so that when scheduling tasks, resources can be reasonably arranged according to the complexity of data objects, improving the efficiency and accuracy of task processing. For compliant data objects that contain multiple data types and have complex relationships between data, this solution can effectively quantify the complexity and is suitable for complex data processing scenarios.

[0034] On the other hand, the number of data types and key data fields in compliant data objects may continue to increase. The complexity calculation method of this solution is based on the independent analysis of each key data field and its association relationship. When adding a new data type or field, it is only necessary to extract and calculate the relevant association value according to the established rules, and then substitute it into the formula for calculation. There is no need to make large-scale changes to the overall architecture, which can well adapt to the growth of data scale and complexity. When a new data type or new association relationship appears, since the solution calculates the complexity by presetting the association value, it is possible to flexibly add the preset value configuration for the new association mode in the system, and then incorporate the new association relationship into the complexity calculation, ensuring the scalability of the technical solution in the face of business changes. In the process of calculating the complexity coefficient, the relevant calculations for each key data field (such as calculating the statistical characteristics of the association value, the number of occurrences, etc.) can be processed in parallel. This allows the use of multi-core processors or distributed computing environments to calculate multiple key data fields at the same time when processing large-scale data objects, greatly shortening the overall calculation time and improving the efficiency of business complexity assessment. The retrieval and use of correlation values ​​in the solution are based on preset configurations. In the actual calculation process, it is only necessary to obtain the corresponding values ​​according to the rules for calculation, which avoids repeated calculations and unnecessary data processing, reduces the waste of computing resources, and further improves processing efficiency. Since the complexity assessment can generate accurate task scheduling weight parameters, the system can reasonably allocate computing resources according to the complexity of different tasks. For tasks with lower complexity, fewer computing resources can be allocated to complete the processing; for complex tasks, more resources are allocated to ensure efficient completion of the tasks. This can avoid over-allocation or under-allocation of resources and improve the overall utilization of computing resources.

[0035] In summary, this technical solution realizes refined task scheduling and efficient processing of compliant data objects through quantitative calculation and generation of weight parameters, improves the efficiency of data processing and compliance management level, and provides strong support for enterprise data management and business development.

[0036] Specifically, the risk coefficient corresponding to the compliant data object and the quantitative value of the timeliness priority are combined with the second complex coefficient to generate the task scheduling weight parameter, including: Identifying risk events contained in the compliance data object; Performing risk probability analysis on the risk events to obtain the corresponding probability of occurrence of each risk event; The risk factor corresponding to each compliant data object is obtained by performing weighted averaging using the probability of occurrence of the risk event and the weight parameter corresponding to the risk event; Identify key time nodes in the processing of compliant data objects and obtain the time difference between each key time node and the current time; Identify the occurrence rate of risk events corresponding to the key time nodes; The quantified value of the timeliness priority is obtained by using the occurrence rate of the risk events corresponding to the key time nodes and the time difference between each key time node and the current time; The quantified value of the timeliness priority is obtained by the following formula: Among them, F represents the quantitative value of timeliness priority; m represents the number of key time nodes; P i represents the occurrence rate of risk events corresponding to the i-th key time node; T i represents the time difference between the i-th key time node and the current time; T c Indicates the preset time difference reference value; Retrieve the second complex coefficient; The risk coefficient corresponding to the compliant data object and the quantitative value of the timeliness priority are combined with the second complex coefficient to generate an initial task scheduling weight parameter; wherein the initial task scheduling weight parameter is obtained by the following formula: Among them, W ctask represents the initial task scheduling weight parameter; k represents the number of data types contained in the compliant data object; R 02i represents the second complexity coefficient corresponding to the i-th data type; F represents the quantitative value of the timeliness priority; U represents the risk coefficient corresponding to the compliance data object; The initial task scheduling weight parameter corresponding to each compliant data object is standardized to generate the task scheduling weight parameter corresponding to each compliant data object.

[0037] The technical effects of the above technical solution are as follows: first, identify the risk events in the compliant data objects, analyze the probability of occurrence of each risk event, and then use the probability of occurrence of the risk event and its corresponding weight parameter for weighted average to obtain the risk coefficient corresponding to each compliant data object, so as to measure the degree of risk faced by the data object. Identify the key time nodes in the data object processing process, calculate the time difference between each key time node and the current time, and determine the risk event occurrence rate corresponding to the key time node. Calculate the quantitative value of the timeliness priority through a specific formula (comprehensively considering the number of key time nodes, the risk event occurrence rate of each node, the time difference and the preset time difference reference value), reflecting the urgency of data processing in terms of time and risk. Retrieve the second complex coefficient calculated previously, combine the risk coefficient and the quantitative value of the timeliness priority, and generate the initial task scheduling weight parameter according to the corresponding formula (considering the number of data types, the second complex coefficient of each data type, etc.), and preliminarily determine the weight of the data object task scheduling. Standardize the initial task scheduling weight parameter to eliminate the dimensional influence between different data, and finally obtain the task scheduling weight parameter corresponding to each compliant data object, which is used to guide task scheduling.

[0038] The solution can generate a comprehensive and accurate initial task scheduling weight parameter by comprehensively considering the risk factor, time priority and second complexity coefficient of the compliant data object. This parameter not only reflects the urgency and importance of data processing, but also takes into account the complexity and processing difficulty of the data itself, thus providing a scientific basis for task scheduling. Through this parameter, the system can automatically prioritize compliant data objects to ensure that critical and urgent tasks are given priority. The identification and analysis of risk events and the calculation of risk factors in the solution help enterprises better identify and manage potential risks. By quantifying the probability of risk occurrence and the corresponding weight, enterprises can more intuitively understand the risk level of each compliant data object, and then take corresponding risk management measures to reduce the possibility and impact of risk occurrence. By identifying key time nodes and calculating the time difference, and combining the occurrence rate of risk events to obtain the quantitative value of time priority, the solution fully considers the timeliness requirements of data processing. This helps enterprises ensure that key data is processed within the specified time and avoid problems caused by delays. At the same time, by optimizing task scheduling, enterprises can improve the efficiency of data processing and reduce unnecessary waiting and delays. By standardizing the initial task scheduling weight parameters, the solution ensures that the task scheduling weight parameters corresponding to each compliant data object are comparable and consistent. This helps enterprises share data and collaborate between different departments or teams, improving overall work efficiency and synergy.

[0039] At the same time, multi-dimensional comprehensive calculations are performed to obtain data from multiple angles such as risk events and time nodes, and the risk coefficient, time priority quantitative value and second complex coefficient are calculated respectively, and then integrated so that the weight parameters can fully reflect the business characteristics and processing requirements of the data object, reduce errors caused by one-sided evaluation, and improve the accuracy of the weight parameters in guiding the actual task scheduling. Operations such as risk event identification and key time node determination have a certain degree of independence. When new risk event types or key time nodes are added due to business changes, only relevant data and calculations need to be supplemented according to established rules, without the need to significantly adjust the overall architecture. It can adapt to changes in data and processing requirements brought about by business development, and ensure the scalability of the weight parameter acquisition solution.

[0040] In summary, this technical solution achieves scientific evaluation and optimization of task scheduling by comprehensively considering the risk factor, timeliness priority and second complexity factor of compliant data objects. This helps enterprises improve the efficiency and accuracy of data processing, reduce risk levels, and enhance overall business operation capabilities and competitiveness.

[0041] See also Figure 3 , a document information entry system based on big data processing, applying the above-mentioned document information entry method based on big data processing, including: Heterogeneous data source parsing module, used for: Extract original document data stored in various media carriers, parse document data in different formats, and generate standardized data input streams; Intelligent data mapping module for: Receive the parsed original document data, align the data according to the template framework, and generate a structured document data object; Risk Signature Screening Module for: Scan the document data objects for risk characteristics, analyze the characteristics of the document data objects, extract key information elements, and identify possible sensitive elements based on the key information elements; Multi-dimensional compliance verification module, used for: Extract business constraints of each verification dimension from the dynamic rule knowledge base and identify the rule version status, and perform multi-dimensional compliance verification on document data objects based on business constraints; Intelligent verification process module for: Based on the review rule set trained on big data, multi-level verification is carried out on compliant data objects, abnormal features are extracted, rule constraint verification is performed on compliant data objects, and credibility scores of compliant data objects are calculated; Audit task scheduling module, used to: Quantify the business complexity, risk factor and timeliness priority of compliant data objects, generate task scheduling weight parameters, and match compliant data objects to the optimal audit path based on the task scheduling weight parameters; Assign audit tasks to audit terminals, and archive compliant data objects to the business database after the audit is completed.

[0042] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A document information entry method based on big data processing, characterized in that: include: Collecting original document data and identifying compliant data objects in the original document data, generating task scheduling weight parameters for the compliant data objects, wherein the process of determining the task scheduling weight parameters includes: retrieving a correlation value based on a data type and a key data field of the compliant data object, obtaining a first complexity coefficient and a second complexity coefficient based on the correlation value, and generating the task scheduling weight parameters in combination with a risk coefficient and a timeliness priority quantified value; Based on the task scheduling weight parameters, each compliant data object is matched to the optimal audit path. After the compliant data object is audited through the optimal audit path, the final processing result is output and archived to the business database.

2. The document information entry method based on big data processing according to claim 1, characterized in that: include: Extract original document data stored in various media carriers, align the original document data to the template framework and generate structured document data objects. The various media carriers include database clusters, optical character recognition files, electronic communication attachment documents and IoT device record streams.

3. The document information entry method based on big data processing according to claim 1, characterized in that: include: Based on big data mining and combined with knowledge graphs, risk feature scanning of document data objects is performed to identify sensitive elements in the preset risk feature library; Based on big data mining technology, feature analysis is performed on document data objects to extract key information elements from the document. The extracted key information elements are compared with the preset risk feature library to identify possible sensitive elements. If sensitive elements are detected, the process will be blocked, and the risk category and impact scope will be determined based on association rule analysis. Real-time warnings will be issued and multi-channel risk notifications will be pushed to the monitoring terminal.

4. The document information entry method based on big data processing according to claim 1, characterized in that: include: Perform multi-dimensional compliance verification on document data objects. If any dimension verification fails, the process will be blocked and a visual diagnostic report will be generated to assist operators in locating the root cause of the anomaly. If the verification succeeds, the compliant data object will be output. Build a dynamic rule knowledge base, extract business constraints of each verification dimension from the dynamic rule knowledge base, and identify the rule version status, which includes existing rules and newly added rules; If it is an existing rule, the document data object is pattern matched with the business constraints of the corresponding dimension; if the business constraints are met, it is determined to be compliant; if it is a new rule, the compatibility of the logical expression of the business constraints and the input parameters is checked, and if there is a compatibility conflict, the conflicting business constraints are frozen; When it is detected that the business constraint is a newly added rule, check whether the business constraint covers all data dimensions. If there are missing dimensions, suspend the business constraint from taking effect.

5. The document information entry method based on big data processing according to claim 1, characterized in that: include: Implement multi-level verification of compliant data objects based on a review rule set trained on big data, identify abnormal patterns and generate credibility scores; Based on the review rule set trained with big data, pattern matching is performed on compliant data objects to extract abnormal features. In combination with a multi-level verification mechanism, rule constraint verification is performed on compliant data objects, the credibility score of compliant data objects is calculated, and the judgment criteria for passing or failing the verification are set based on the preset score threshold; If the verification fails, the process will be blocked and an abnormal pattern analysis report will be generated. If the verification passes, the compliant data object will be output to the next stage.

6. The document information entry method based on big data processing according to claim 1, characterized in that: include: Intelligent task routing is performed on compliant data objects that have passed verification, and each compliant data object is matched to the optimal audit path based on business complexity, risk factor, and timeliness priority. Quantify the business complexity, risk factor and timeliness priority of compliant data objects, generate task scheduling weight parameters, and match compliant data objects to the optimal audit path based on the task scheduling weight parameters; Generate audit tasks based on compliance data objects, assign audit tasks to compatible audit terminals, and after the audit is completed, output the final processing results and archive them in the business database.

7. The document information entry method based on big data processing according to claim 6, characterized in that: The quantitative calculation of the business complexity, risk factor and timeliness priority of the compliant data object to generate the task scheduling weight parameter includes: Extracting the data types contained in the compliant data object and the key data fields contained in each data type; Retrieving a preset correlation value between the key data field contained in each data type and other key data fields of the same data type in the data type; Obtaining a first complexity coefficient corresponding to each data type by using a preset correlation value between a key data field contained in each data type and other key data fields of the same data type within the data type in which the key data field is located; Retrieve the preset correlation value between the key data fields contained in each data type and other data types; Obtaining a second complexity coefficient corresponding to each data type by using the first complexity coefficient corresponding to each data type in combination with a preset correlation value between a key data field contained in each data type and other data types; The risk coefficient corresponding to the compliant data object and the quantified value of the timeliness priority are combined with the second complex coefficient to generate a task scheduling weight parameter.

8. The document information entry method based on big data processing according to claim 7, characterized in that: The task scheduling weight parameter is generated by using the risk coefficient corresponding to the compliant data object and the quantified value of the timeliness priority in combination with the second complex coefficient, including: Identifying risk events contained in the compliance data object; Performing risk probability analysis on the risk events to obtain the corresponding probability of occurrence of each risk event; The risk factor corresponding to each compliant data object is obtained by performing weighted averaging using the probability of occurrence of the risk event and the weight parameter corresponding to the risk event; Identify key time nodes in the processing of compliant data objects and obtain the time difference between each key time node and the current time; Identify the occurrence rate of risk events corresponding to the key time nodes; The quantified value of the timeliness priority is obtained by using the occurrence rate of the risk events corresponding to the key time nodes and the time difference between each key time node and the current time; The quantified value of the timeliness priority is obtained by the following formula: Among them, F represents the quantitative value of timeliness priority; m represents the number of key time nodes; P i represents the occurrence rate of risk events corresponding to the i-th key time node; T i represents the time difference between the i-th key time node and the current time; T c Indicates the preset time difference reference value.

9. The document information entry method based on big data processing according to claim 8, characterized in that: The task scheduling weight parameter is generated by using the risk coefficient corresponding to the compliant data object and the quantified value of the timeliness priority in combination with the second complex coefficient, and further comprising: Retrieve the second complex coefficient; The risk coefficient corresponding to the compliant data object and the quantitative value of the timeliness priority are combined with the second complex coefficient to generate an initial task scheduling weight parameter; wherein the initial task scheduling weight parameter is obtained by the following formula: Among them, W ctask represents the initial task scheduling weight parameter; k represents the number of data types contained in the compliant data object; R 02i represents the second complexity coefficient corresponding to the i-th data type; F represents the quantitative value of the timeliness priority; U represents the risk coefficient corresponding to the compliance data object; The initial task scheduling weight parameter corresponding to each compliant data object is standardized to generate the task scheduling weight parameter corresponding to each compliant data object.

10. A document information entry system based on big data processing, applied in a document information entry method based on big data processing according to any one of claims 1 to 9, characterized in that: include: Heterogeneous data source parsing module, used for: Extract original document data stored in various media carriers, parse document data in different formats, and generate standardized data input streams; Intelligent data mapping module for: Receive the parsed original document data, align the data according to the template framework, and generate a structured document data object; Risk Signature Screening Module for: Scan the document data objects for risk characteristics, analyze the characteristics of the document data objects, extract key information elements, and identify possible sensitive elements based on the key information elements; Multi-dimensional compliance verification module, used for: Extract business constraints of each verification dimension from the dynamic rule knowledge base and identify the rule version status, and perform multi-dimensional compliance verification on document data objects based on business constraints; Intelligent verification process module for: Based on the review rule set trained on big data, multi-level verification is carried out on compliant data objects, abnormal features are extracted, rule constraint verification is performed on compliant data objects, and credibility scores of compliant data objects are calculated; Audit task scheduling module, used to: Quantify the business complexity, risk factor and time priority of compliance data objects, generate task scheduling weight parameters, match compliance data objects to the optimal audit path based on the task scheduling weight parameters, assign audit tasks to audit terminals, and archive compliance data objects to the business database after the audit is completed.

Citation Information

Patent Citations

  • Method and device for processing refund data

    CN107909362A

  • Receipt information input method and device based on data processing, and storage medium

    CN109739957A

  • Information processing system based on RPA robot and working method thereof

    CN116340453A

  • Method applied to enterprise OA collaborative workflow optimization

    CN118761745A

  • Compliance inspection system and method for enterprise business process

    CN119378993A

Cited By

  • Customs declaration intelligent dispatching and cooperative processing system

    CN120471403A

  • Financial bill intelligent identification and verification method based on machine learning

    CN120612190A

  • Machine learning-based intelligent identification and verification method for financial bills

    CN120612190B

  • Intelligent checking system and method for customs affair compliance

    CN120672088A

  • Receipt interface type generation method and system based on target data

    CN120687090A